Data Engineer (Data Lakehouse)
Comm It — Poland · Posted ~4 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
We’re looking for a Middle Data Engineer to take ownership of our data lake — the system of record for millions of financial events every day, including bets, wallet transactions, and live odds, serving 12M+ active users.You’ll be responsible for designing and maintaining reliable data pipelines and defining how data is ingested, stored, retained, reconciled, and governed across AWS S3 and Snowflake/Databricks.
Your work will ensure that Analytics, Finance, and Regulatory teams have access to accurate, consistent, and fully traceable data — with zero drift from source systems.
Location: Kraków, Poland.
Hybrid — 2 days per week from the office.
What you will do:
Own the lakehouse architecture: bronze/silver/gold layers, Iceberg/Delta tables, schema evolution.Land operational data via CDC streaming (Kafka, Debezium), handling late and duplicate events.Design data layout for speed and cost: partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.Own retention and archival: storage tiering, regulatory retention, immutability, GDPR deletion.Guarantee correctness: freshness SLAs, drift detection, reconciliation against the source wallet and ledger systems.Own governance: catalog and lineage, row/column access control, PII masking, encryption, audit trails.Monitor ingestion health, data anomalies, and cloud storage/compute spend.
Requirements:
Must-have:
3+ years hands-on in a production lakehouse environment.Lakehouse architecture — bronze/silver/gold layering, an open table format (Iceberg, Delta, or Hudi), schema evolution.Data layout & query optimization at TB+ scale — partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.Cloud lakehouse/DWH in production — Snowflake, Databricks, or BigQuery.CDC & streaming ingestion — Kafka + Debezium or equivalent; late, duplicate and out-of-order events.Strong SQL and data modeling — enough relational grounding to reason about the OLTP systems you capture from.
Critical for financial ledgers.Correctness — freshness SLAs, drift detection, reconciliation against source wallet/ledger systems.Governance — catalogs, lineage, row/column access control, PII masking, retention, GDPR deletion.Cloud object storage — S3 or GCS, plus storage tiering and archival.Python and an orchestrator — Airflow or Dagster, as tools.
Nice to have:
Fintech, iGaming, or another regulated, audit-heavy environment.Cost monitoring / FinOps for storage and compute spend.Hudi specifically; Dagster specifically.Immutability / WORM regulatory retention.
We have 142,744 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 142,744 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume