Summary
✨ AI‑Generated
A senior data engineering opportunity for a solution-oriented engineer who can design and optimize scalable batch and near-real-time integration pipelines. The role emphasizes Python, PySpark, SQL, cloud-native architectures, workflow orchestration, automation, reliability, and high data quality while bridging source systems with analytical environments.
Highlights
Senior technical role focused on modern cloud-native data engineering, scalable pipelines, workflow automation, reliability, and data quality, with close collaboration across engineering and architecture teams.
Description
We are seeking a highly skilled, solution-oriented Senior Data Integration Engineer with deep expertise in modern data engineering, cloud-native architectures, and robust pipeline development.
In this role, you will be a senior technical driver in creating, optimizing, and modernizing advanced data integration systems.
You will collaborate closely with product, engineering, and architecture teams to bridge raw source environments and analytical databases.
If you excel in cloud ecosystems, thrive on automating complex data workflows, and value precision, reliability, and data quality above all, we encourage you to apply!
Responsibilities
Pipeline Architecture & Development: Design, build, and optimize scalable, reliable batch and near-real-time ETL/ELT pipelines using Python, PySpark, SQL, and modern cloud integration enginesOrchestration & Automation: Develop and manage complex workflow orchestrations (using Apache Airflow or cloud native schedulers) and automate ingestion routines to minimize manual operationsData Modeling & Warehousing: Design and implement modern data warehouse/lakehouse layers (using Snowflake, ClickHouse, Azure Synapse, or Redshift), establishing optimal partitioning, indexing, and Slowly Changing Dimension (SCD Type 2) patternsData Quality & Testing Integration: Establish rigorous data quality checks and validation frameworks utilizing tools like dbt (data build tool), Soda, or customized PySpark testing suitesCollaboration & Design: Work closely with product owners, business analysts, and systems architects to define data requirements, analyze technical constraints, design Source-to-Target Mappings (STTM), and make critical architectural decisionsCode Quality & DevOps: Maintain a clean, modular code repository.
Lead code reviews, enforce engineering standards, and configure robust CI/CD pipelines (Azure DevOps, GitLab CI, or GitHub Actions) with Docker containersTechnical Documentation: Deliver comprehensive, clear technical specs, metadata lineage documentation, architectural diagrams, and data dictionaries
Requirements
Experience: 5+ years of hands-on experience in data engineering, data warehousing, database design, and end-to-end data integrationETL & Integration Tools: Advanced knowledge of Cloud Integration tools such as Azure Data Factory (ADF), AWS Glue, or GCP DataflowOrchestration & Real-Time Ingestion: Proficiency in workflow orchestrators like Apache Airflow and exposure to CDC (Change Data Capture) or real-time streaming tools (e.g., Kafka, Debezium)Core Technical Stack: Strong production-level coding skills in SQL (advanced optimization/stored procedures), Python, and PySpark / Apache SparkAnalytical Databases & Cloud Warehouses: Experience working with high-performance databases and cloud-native systems (e.g., Snowflake, ClickHouse, PostgreSQL, MS SQL Server, or Azure Synapse)Methodologies: Master-level understanding of data modeling practices (OLAP, OLTP, Star/Snowflake schemas, Delta Lake/Lakehouse patterns, and Data staging processes)DevOps & CI/CD: Hands-on experience with version control (Git) and building automated deployment pipelines (CI/CD) for data productsCommunication & English: Proven ability to articulate complex technical ideas clearly to both business stakeholders and developers.
Fluency in English (Upper-Intermediate level or higher)
Nice to have
Data Transformation & Quality Tools: Deep knowledge of dbt (data build tool) and schema validation practicesContainerization: Experience using Docker or Kubernetes to package and deploy data applicationsServerless Engineering: Experience building lightweight, serverless ingestion services (e.g., using AWS Lambda / Azure Functions and RESTful APIs)
We offer
We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageSix trust days per year (sick leave without a medical certificate)Coverage of psychology sessions of your choiceDiscounts for fitness clubs and sports programsBenefits package (sports activities, a variety of stores and services)
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups.
With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.
Experience the freedom of remote work from anywhere in Kyrgyzstan, whether it's the comfort of your home or our modern office in Bishkek.