Summary
✨ AI‑Generated
A senior data platform engineering role focused on building reliable, scalable infrastructure for massive datasets, distributed workloads, and advanced AI systems.
Highlights
Work on advanced large-scale data infrastructure supporting cutting-edge research and high-performance computing workloads.
Description
Senior Data Platform Engineer – Distributed Data Processing / AI Infrastructure / GPU
We are currently partnered with a highly technical AI research organisation developing infrastructure to support next-generation frontier AI models.
This role focuses on architecting, building, scaling, and maintaining large-scale data infrastructure, with research teams as the primary customers of the data platform.
You will work across petabyte-scale data processing, distributed workloads, GPU-accelerated processing, reproducibility, lineage, and data platform reliability.
Key responsibilities
Scale and automate data processing infrastructure to handle petabytes of data while ensuring reliable day-to-day operation.Design execution layers for pipeline stages with different computational requirements, selecting appropriate processing engines and maintaining efficient end-to-end throughput.Optimise compute resource utilisation across large-scale workloads, including GPU resources for compute-intensive data processing.Build reliable pipelines with graceful failure recovery, restartability, and comprehensive observability.Identify and diagnose performance bottlenecks and attribute compute usage and costs to individual pipeline stages.Establish systems for dataset versioning, lineage, reproducibility, and traceability across the data processing lifecycle.Ensure intermediate outputs and final datasets can be reliably reproduced for specific and evolving research requirements.
Key requirements
Bachelor’s degree in a relevant field such as Computer Science, Computer Engineering, or Software Engineering.5+ years of experience designing, implementing, and managing large-scale distributed data processing systems or large-scale distributed ML data frameworks.Recent experience with technologies such as Ray, Apache Spark, workflow orchestrators, Apache Arrow, and/or Parquet.Demonstrated ownership of a data processing system operating under real throughput, reliability, and cost pressures.Experience reasoning about large-scale pipeline failures and diagnosing why processing pipelines break.
Keywords
Senior Data Platform Engineer / Data Platform / Data Infrastructure / Frontier AI / AI Models / Large-Scale Data Processing / Petabyte-Scale Data / Distributed Data Processing / Distributed ML Data Frameworks / Data Processing Pipelines / Pipeline Execution / Processing Engines / Compute Optimisation / GPU-Accelerated Data Processing / GPU Workloads / Throughput Optimisation / Cost Optimisation / Performance Bottlenecks / Observability / Failure Recovery / Restartability / Dataset Versioning / Data Lineage / Reproducibility / Traceability / Data Governance / Data Specifications / Model Training Pipelines / Ray / Apache Spark / Workflow Orchestrators / Apache Arrow / Parquet / Kubernetes / Slurm / Docker / Enroot / Vector Databases / Heterogeneous Workloads / Distributed Systems / Data Processing Infrastructure / AI Research Infrastructure / Cross-functional Collaboration
By applying to this role, you understand that we may collect your personal data and process it in line with our privacy policy.
https://eu-recruit.com/wp-content/uploads/2025/01/European-Tech-Recruit-Privacy-Notice-2025.pdf