Senior Data Engineer

Undisclosed 7 — United Kingdom · Posted ~3 hours ago

Senior Contract Hybrid

Skills

Data engineering Data pipelines Cloud technologies Machine learning infrastructure Artificial intelligence GCP Python Cloud Machine Learning AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior data engineering role focused on designing and delivering scalable data products. You will own data pipelines, work with complex sources, and help build AI-driven solutions using modern cloud platforms.

Highlights

Work on advanced data and AI initiatives with opportunities to build large-scale solutions using modern engineering practices and cloud technologies.

Description

Role Title: Senior Data Engineer Duration: 12 months with possibility of extension Location: London. Hybrid, 2 days per week onsite Rate: TBC Role purpose / summary We see a world in which advanced applications of machine learning (ML) and artificial intelligence (AI) will allow us to develop novel therapies for existing diseases and to respond quickly to emerging or changing diseases with personalised drugs, driving better outcomes at reduced cost with fewer side effects. It is an ambitious vision and delivering it will require products and solutions at the cutting edge of Machine Learning and AI. We’re looking for a senior data engineer (contractor) to help us make this vision a reality. Strong candidates will have a track record of shipping data products derived from complex sources and will have owned that process from initial pipeline design through to production scale. We have a commitment to quality, so successful candidates will be able to use modern cloud tooling and techniques to deliver reliable data pipelines and continuously improve them. This role calls for in interest in solving hard problems in Artificial Intelligence and Machine Learning, and for doing that work collaboratively as part of a team. The successful candidate will make substantial use of AI agents to develop software in this role. Key responsibilities Build, operate and then automate the pipelines behind our medical imaging data and combine that imaging data (DICOM data) with other data modalities including tabular data.Design data storage, transfer and transformation utilities that make it fast to move and harmonise multi-terabyte image datasets on Google Cloud Platform (GCP) by developing standardised onboarding processes with imaging sites and vendors. This includes designing secure inbound and outbound exchange with automated data transfer and Quality Control.Handle derived imaging artefacts. Link segmentations and annotations (for example RTSTRUCT, NIfTI) back to their source series and deliver analysis-ready dataset snapshots into the imaging analysis platforms and ML environments for science teams to work with.Automate the path for curating and harmonising imaging datasets. Parse and normalise metadata, validate incoming manifests against the agreed metadata standard, reconcile images against clinical/tabular and other biomarker data, and build quality and de-identification checks to replace manual review.Write production-grade Python: tested, reviewed, instrumented and documented well enough that the team can run it long-term.Work with ML engineers and software engineers to shape datasets around what the models and the science need.Work with imaging leads to apply ML and AI techniques to ingestion, including by doing anomaly detection over metadata and series structure, using automated detection of missing, duplicate or mismatched studies, applying classifiers for burned-in pixel PHI, and building agent-driven triage of failed loads. Essential qualifications Strong Python for data engineering, including data transformation, job monitoring and schema management.Experience handling large binary/blob datasets in object storage at scale (Google Cloud Storage, Azure Blob Storage/ADLS, Amazon S3 or equivalent) in a cloud environment.Significant SQL experience, including schema design.One of: a PhD in a computational discipline, MSc in computational discipline + 2 years’ experience with imaging data (from any domain), or 2 years’ hands-on experience with medical imaging data (such as DICOM data).Experience building and running ETL/ELT pipelines with an orchestration framework.Experience with CI/CD, agile software development and DevOps.Ability to work with ambiguous requirements and decide on the approach independently. Desirable qualifications Hands-on experience with DICOM data; metadata and tags, multi-frame and series structure, de-identification including burned-in pixel PHI, and the standard Python tooling (pydicom, SimpleITK, dcm2niix, DCMTK).Deep, practical BigQuery and/or SQL: schema design, partitioning and clustering, query and cost optimisation on large tables.Agentic engineering; using coding agents as a day-to-day part of how you build and building agent-driven data workflows.Google Cloud; services such as Cloud Run, GKE, Artifact Registry and Cloud SQL.Modern columnar and array formats (Parquet, Arrow, Zarr) and thoughtful storage layout for large datasets.Infrastructure as Code (IaC); ideally Terraform, Docker containers.Experience working with sensitive data under GDPR, HIPAA or clinical trial data governance.Experience building secure, audited data exchange with external parties like imaging vendors, academic collaborators and Contract Research Organisations, including staging containers, transfer controls and provenance tracking.Background in biology, medicine or biomedical data (genomics, transcriptomics, proteomics, EHR, clinical images). Key Skills Python, GCP, DICOM, Parquet, Image Data, Cloud, SQL, BigQuery All profiles will be reviewed against the required skills and experience. Due to the high number of applications we will only be able to respond to successful applicants in the first instance. We thank you for your interest and the time taken to apply!