Senior Distributed Data Engineer

European Tech Recruit — Germany · Posted ~2 hours ago

Senior Full-time Hybrid Visa History ✓

Skills

Distributed systems Data infrastructure Data processing Scalable systems Cloud infrastructure Data platforms AI infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior data engineering role building and operating distributed platforms that process massive datasets. The role focuses on reliability, scalability, performance optimization, and enabling advanced data-driven applications.

Highlights

Work on large-scale data infrastructure supporting advanced AI systems, solving complex distributed computing challenges with a highly technical team.

Description

Senior Data Infrastructure Engineer Berlin, Germany | Hybrid - 3 days onsite weekly Join a highly technical AI research and engineering organisation building the data infrastructure required to support next-generation AI models at massive scale. As a Senior Data Infrastructure Engineer, you will architect, build and operate the distributed systems that process petabytes of data and connect large-scale data workflows with cutting-edge AI research. The Senior Data Infrastructure Engineer will focus on scalability, reliability, throughput, resource utilisation and cost efficiency, ensuring that complex data-processing workloads can run consistently across large distributed environments. This Senior Data Infrastructure Engineer role is ideal for someone who enjoys solving difficult distributed-systems problems and has experience owning data platforms under real-world throughput, reliability and cost pressures. Key Responsibilities Design, build and scale distributed data-processing infrastructure capable of handling petabyte-scale workloads.Develop execution layers for pipeline stages with different computational requirements, selecting appropriate processing engines and technologies.Optimise compute utilisation across large-scale workloads, including GPU-accelerated data-processing stages.Build highly reliable pipelines with graceful failure recovery, restart-ability, monitoring and comprehensive observability.Profile large-scale workloads to identify throughput bottlenecks and optimise compute utilisation and infrastructure costs.Establish robust systems for dataset versioning, lineage, reproducibility and traceability across the data-processing lifecycle.Ensure intermediate and final datasets can be reliably reproduced against specific research requirements.Build and maintain data infrastructure that integrates effectively with large-scale machine learning training pipelines.Evaluate emerging technologies and architectural approaches to improve scalability, reliability, performance and cost efficiency.Establish technical documentation, standards and best practices across the data platform. Required Experience & Skills Bachelor's degree or higher in Computer Science, Computer Engineering, Software Engineering or a related discipline.At least 5+ years of full-time experience designing, implementing and operating large-scale distributed data-processing systems or ML data infrastructure.Demonstrated ownership of production data-processing systems operating under significant throughput, reliability and cost pressures.Strong experience with distributed processing frameworks such as Ray, Apache Spark or Flink.Experience with workflow orchestration and large-scale pipeline execution.Strong understanding of performance profiling, throughput optimisation and distributed computing.Experience working with Apache Arrow, Parquet or comparable high-performance data formats.Experience with dataset versioning, lineage, reproducibility and traceability.Experience with workload managers or schedulers such as Ray, Kubernetes or Slurm.Familiarity with containerisation technologies including Docker or Enroot.Experience designing systems for fault tolerance, observability and failure recovery.Strong software engineering skills and the ability to make sustainable technical decisions in rapidly evolving environments.Desired: Experience supporting GPU-accelerated data-processing workloads; Experience working with cloud infrastructure and large-scale compute environments; Familiarity with vector databases and modern data infrastructure; Experience with infrastructure-as-code; Experience working directly with ML training infrastructure or frontier AI workloads. Why Apply? Build the core data infrastructure supporting large-scale AI research and model development.Work at petabyte scale, solving complex distributed-systems and data-platform challenges.Have genuine architectural ownership over how large-scale workloads are executed, monitored and optimised.Work closely with highly technical Research and Engineering teams.Evaluate and introduce new technologies rather than simply maintaining an established data platform.Solve problems spanning distributed computing, GPU utilisation, reliability, observability, throughput and infrastructure cost. Apply now or send a copy of your CV, referencing the title and location, and with a short intro to cw@eu-recruit.com. By applying to this role you understand that we may collect your personal data and store and process it on our systems. For more information please see our Privacy Notice (https://eu-recruit.com/about-us/privacy-notice/)