Summary
✨ AI‑Generated
A data engineering role focused on designing pipelines and infrastructure that process large volumes of structured information and support intelligent digital products.
Highlights
Build large-scale data systems, work on impactful technology, and contribute to reliable information platforms used by professionals.
Description
Data Engineer
Amsterdam | In office | Full-time | 1-3 years experience
About Moonlit
Moonlit is the legal publisher of today.
From our headquarters in Amsterdam we serve tens of thousands of legal professionals globally, either through our platform, our API or our MCP server.
Behind it all sits Europe's largest legal database, enriched and structured to power semantic search, AI assistants and the legal AI products built on top of us.
Our mission
Make legal information more efficient, transparent and accessible.
For everyone.
The role
As a Data Engineer at Moonlit you'll build the data layer that powers a platform used daily by lawyers, judges and policymakers.
Our users work at the highest levels and demand quality.
You'll design and implement the pipelines and infrastructure that bring millions of legal documents into our ecosystem, keep them accurate and up to date and make them fast to search.
We build enterprise and government grade software.
What you will do
Build and optimize ETL pipelines using Python, PySpark and DatabricksAdd new legal data sources: scraping, parsing and normalizing documents in formats like HTML, XML and PDFEnrich and structure existing datasets (metadata, references between documents, versioning of legislation)Maintain and optimize our Elasticsearch indexes for fast and reliable searchMonitor data quality, freshness and coverage across all sourcesImprove performance, reliability and cost of existing pipelinesManage data infrastructure on AzureCollaborate with platform developers, AI engineers, legal experts and business stakeholders
Data ownership
Own end-to-end data workflows: from source to searchable, enriched documentBuild for production: tested, monitored and scalableBe the go-to person for how our data is structured and where it comes from
Our tech stack
Data processing: Python, PySpark, Databricks
Search & storage: Elasticsearch, Turbopuffer, Azure SQL
Cloud: Azure, AWS, GCP
Who are you?
1-3 years of experience in data engineering or a related fieldStrong Python skills with experience in PySpark and large-scale data processingSolid SQL and a good understanding of data modelingHands-on experience with at least one major cloud platform (Azure, AWS or GCP)Experience building ETL pipelines and data transformation workflowsComfortable with messy, unstructured dataExperience with Databricks and/or Elasticsearch is a plusA pragmatic mindset: you choose robust solutions over clever ones
Why Moonlit?
At Moonlit your work has real-world impact.
The data systems you build support tens of thousands of legal professionals globally and have the potential to help millions gain better access to justice.
We're a quality-first, engineering-driven team that chooses scalable, modern and elegant technologies.
You'll have room to take initiative, shape how we handle data and grow personally and professionally in a fast-paced scale-up.
You'll work with talented developers and legal experts in an open culture.
What we offer
Competitive salary with equity packageA very nice office with a big gardenAn international, mission-driven work environment