Summary
✨ AI‑Generated
An AWS-focused data engineering role responsible for building and maintaining end-to-end batch and near-real-time pipelines, implementing ETL/ELT workflows, and developing reliable lake and warehouse layers. The position combines SQL, Python and PySpark with cloud services, orchestration, data quality controls, CI/CD, testing, monitoring, and collaboration with analytics and machine learning teams.
Highlights
End-to-end data engineering role covering batch and near-real-time pipelines, cloud data platforms, data quality, monitoring, analytics enablement, and modern DataOps/DevOps practices.
Description
Key ResponsibilitiesDesign, develop, and maintain end-to-end data pipelines (batch and near real-time) on AWS Data PlatformBuild and manage ETL/ELT workflows using AWS services (e.g., AWS Glue, S3, Redshift, Athena, EMR), dbt and orchestration tools such as AirflowImplement data ingestion patterns from diverse sources (databases, APIs, files, event streams) into lake/warehouse layers such as raw, cleansed, and curated data layersDevelop transformation logic using SQL and Python/PySpark for cleansing, enrichment, and standardisationImplement robust data quality checks, reconciliation controls, and monitoring/alerting for failures and anomaliesCollaborate with data analysts/data scientists to model datasets for analytics and machine learning consumption.Contribute to DataOps/DevOps practices: version control, CI/CD, automated testing, release management, and operational support.Produce and maintain technical documentation (data flows, mappings, job schedules, runbooks, and operational procedures)Optimise Data Pipeline performance and Support workflow orchestration and schedulingSupport production deployments and operationsRequired Skills & Experience8–10 years’ experience as a Data EngineerAdvanced SQL skillsHands-on experience working with Teradata and Siebel CRM data setsExperience delivering data pipelines in a large-scale enterprise data platform environmentStrong hands-on AWS experience with common data services such as : Amazon S3, AWS Glue, Amazon Redshift, Amazon Athena, Amazon EMR and dbtStrong programming capability in Python and strong data transformation experience using PySpark (preferred) and/or Spark.Advanced SQL skills (query optimisation, complex joins, window functions, performance tuning)Experience with workflow orchestration tools such as AirflowSolid understanding of data warehousing concepts (dimensional modelling, partitioning, incremental loads, CDC concepts).Experience implementing monitoring, logging, alerting, and operational support processes.Strong communication skills and ability to work with stakeholders to translate requirements into data deliverablesTelco Industry Experience is highly desirable