Summary
✨ AI‑Generated
A research-focused organization is seeking a senior data platform engineer to design scalable data infrastructure supporting advanced artificial intelligence development.
Highlights
Build foundational data infrastructure for advanced AI research with large-scale distributed systems and engineering challenges.
Description
Join a mission-led, non-profit organisation working to make advanced artificial intelligence safer by design.
We are looking for an accomplished Senior Data Platform Engineer to build the data foundations behind next-generation AI research.
You will own critical infrastructure that processes data at petabyte scale, enabling researchers to create dependable, traceable datasets for frontier model development.
This is an opportunity to work where large-scale distributed systems, machine learning and research engineering meet.
You will treat the platform as a product, with research teams as its users, and translate evolving experimental needs into robust production systems.
What You’ll Do
Architect and automate distributed data pipelines capable of processing petabytes of data.Create execution layers that support pipeline stages with different compute and processing requirements.Select appropriate processing technologies and optimise throughput across the full data lifecycle.Improve the efficiency and cost of heterogeneous workloads, including GPU-intensive processing.Engineer pipelines that recover gracefully from failures and can be safely restarted.Implement comprehensive monitoring and observability across large-scale workloads.Investigate performance constraints and attribute infrastructure usage and costs to individual processing stages.Build dependable systems for dataset versioning, lineage, reproducibility and end-to-end traceability.Ensure intermediate artefacts and final datasets can be recreated as research requirements evolve.Define clear data specifications and maintain documentation that supports internal governance.Work closely with Research, Product and Engineering teams to integrate datasets with model-training workflows.Assess emerging technologies and introduce them where they offer meaningful improvements in performance, reliability, scalability or cost.
What You’ll Bring
A bachelor’s degree in computer science, computer engineering, software engineering or another relevant discipline.At least five years of experience building and operating large-scale distributed data-processing systems or distributed machine-learning data frameworks.Recent hands-on experience with technologies such as Ray, Apache Spark, workflow orchestration tools, Apache Arrow or Parquet.Evidence of owning a data-processing platform under genuine throughput, reliability and cost constraints.A strong understanding of how complex pipelines fail, how to diagnose those failures and how to prevent them from recurring.Experience profiling and improving performance and cost across varied workloads, including GPU-accelerated stages.Practical knowledge of dataset lineage, versioning and reproducibility systems.Experience with workload-management platforms such as Ray, Kubernetes or Slurm.Familiarity with container technologies such as Docker or Enroot.Familiarity with modern data infrastructure, including areas such as vector databases.The ability to work effectively across research and engineering disciplines, establish best practices and communicate technical decisions clearly.
The technology landscape will continue to change as the organisation’s research ambitions grow.
You will have the scope to challenge established approaches, evaluate new possibilities and help shape a sustainable platform for technically demanding AI work.
If you are an experienced data-platform engineer excited by difficult distributed-systems challenges and the opportunity to support safer frontier AI, apply now or email nk@eu-recruit.com