Summary
✨ AI‑Generated
A data engineering position focused on building, maintaining, and troubleshooting cloud-based data pipelines. The role requires experience with distributed processing, automation, and modern cloud services.
Highlights
Hands-on data engineering role working with large-scale cloud platforms, modern data pipelines, and advanced analytics technologies.
Description
Data Engineer
Up to £65k per annum
Apache Spark Python AWS Cloud Data Pipelines
A hands-on data engineering role within a large-scale cloud data programme, responsible for building, maintaining, and troubleshooting data pipelines using Apache Spark, PySpark, Apache Airflow, and a broad suite of AWS services.
You will apply strong analytical and engineering skills to deliver trusted, well-governed data assets in a modern, cloud-native environment.
About Scrumconnect
Scrumconnect is a leading UK technology consultancy delivering digital transformation across public and private sectors, contributing to over 20% of the UK's major citizen-facing public services.
We specialise in cloud engineering, data platforms, and agile delivery, helping clients build scalable, secure, and user-centred digital solutions that create real impact.
Working arrangement:
This role is hybrid.
Candidates must be willing and able to travel to the Newcastle office once per week.
Remaining days may be worked remotely from anywhere in the UK.
About the role
You will work as a Data Engineer on a complex, cloud-based data programme - designing, building, and maintaining data pipelines that process large volumes of data across a modern AWS-native stack.
Using Apache Spark and PySpark for distributed data processing, Apache Airflow for orchestration, and a range of AWS services for storage, compute, and analytics, you will help deliver reliable, well-governed data assets to downstream users.
You will apply strong data analysis skills to identify root causes of data issues, work with dimensional data models and slowly changing dimensions, and implement infrastructure as code using Terraform.
Familiarity with engineering best practices and the ability to translate customer expectations into applied technical functionality are key to success in this role.
Key responsibilities
Data pipeline development
Build and maintain scalable data pipelines using Apache Spark and PySpark, processing and transforming large datasets across distributed cloud infrastructure.
Workflow orchestration
Configure and manage Apache Airflow DAGs for task orchestration, ensuring reliable scheduling, monitoring, and execution of data processing workflows.
Root cause analysis
Perform data analysis to identify and resolve root causes of pipeline failures and data quality issues - including reviewing EMR output logs and CloudWatch metrics.
Data modelling
Apply understanding of dimensional data models and slowly changing dimensions (SCD) to design and maintain well-structured, analytically trusted data assets.
Infrastructure as code
Provision and manage cloud infrastructure using Terraform.
Containerise solutions using Docker and manage deployments through GitLab CI/CD pipelines and release tagging.
Security & encryption
Apply understanding of both Server Side and client-side encryption patterns within AWS.
Work within IAM policies and data governance standards appropriate to a regulated government environment.
Technical skills required
Languages & analytics
Python - primary language for pipeline development and data processingSQL - used for querying, transformation, and validation across data storesPySpark - Power BI for distributed data processing using Apache Spark on AWS EMRFamiliarity with basic data structures for constructing robust, scalable solutionsData processing & orchestration
Apache Spark - understanding of distributed data processing architecture and executionApache Airflow - configuring DAGs and managing task orchestration at scaleJupyter Notebooks - for exploratory data analysis and pipeline prototypingUnderstanding of dimensional data models and slowly changing dimensions (SCD Types 1, 2, 3)Data analysis skills to identify root cause of issues within pipelines and data assetsAWS services
Amazon EMR - running Spark workloads and reviewing output logsAmazon Athena - ad hoc querying of data in S3Amazon Textract and Comprehend - familiarity with AI/ML document extraction and NLP servicesAWS S3, IAM, CloudWatch, EC2, ECR - core platform services used day-to-dayAWS console proficiency - navigating, configuring, and monitoring servicesUnderstanding of Server Side and client-side encryption within AWSInfrastructure, DevOps & delivery
Terraform - Infrastructure as Code for provisioning and managing AWS environmentsDocker - containerisation of data engineering solutionsGitLab - source code management, CI/CD pipeline configuration, release tagging, and component versioningFamiliarity with engineering best practicesAbility to translate customer expectations into applied, functional technical solutionsTechnology stack at a glance
PythonPySparkSQLApache, Power BI SparkApache AirflowJupyter NotebooksDimensional modelling/SCDAWS EMRAmazon AthenaAWS S3AWS IAMAWS CloudWatchAWS EC2/ECRAmazon TextractAmazon ComprehendTerraformDockerGitLab CI/CDGitLab
Diversity & Inclusion
At Scrumconnect Consulting, we are proud to be a Disability Confident employer and strongly committed to building an inclusive workforce that reflects the diverse communities we serve.
We welcome applications from individuals of all backgrounds, identities, and abilities.
Our HR and recruitment processes are fair, accessible, and designed to ensure equal opportunities for all.
If you need any adjustments during the application or interview process, please let us know—we are here to support you.