Site Reliability Engineer - Data Platform

Imc Trading — Netherlands · Posted ~3 hours ago

Senior Full-time Visa History ✓

Skills

SRE Distributed data systems Linux Kubernetes Deployment automation Observability Scalability Bare metal Distributed systems Automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a data platform team as an experienced Site Reliability Engineer and help operate and evolve distributed data systems across bare-metal Linux and Kubernetes. You will standardize deployments, strengthen observability, automate operational workflows, and improve the scalability and reliability of foundational services used by technical and research teams.

Highlights

Work on foundational data infrastructure supporting traders, quantitative researchers, and engineering teams. The role combines distributed systems, Kubernetes, observability, automation, and scalability, with substantial ownership over critical platform services and engineering standards.

Description

IMC operates on the cutting-edge use of technology to create a competitive edge over the competition. We also grow quick and have plenty of complex technical challenges. We're looking for an experienced SRE with strong background in managing distributed data systems on both bare metal Linux and Kubernetes. We want someone who can help us standardize deployments, elevate observability, and improve automation as we scale our data platform and other critical data services. You will join our Data Platform team, part of our local data team that builds and runs the systems that are used by traders, quant researchers and engineering teams, for all their data needs. The Data Platform team is responsible for the foundational platform that our data frameworks and tooling is built on top of. This includes observability, scalability and supporting standardised deployments. Your Core Responsibilities As an SRE within IMC you will join a sub-team that takes a central role in all the data needs and you’ll be working to: Design, implement and operate our data platforms.Improve observability so we catch issues before our users do.Build automation to reduce toil and allow our systems to scaleSupport and own reliability of critical services (e.g. HDFS, Kafka and Dremio)Drive long-term architectural improvements, not just fixing issues, but preventing them. Your Skills And Experience Strong experience managing distributed data platforms (e.g. Kafka, Hadoop, Spark, Dremio); including full installation, debugging and performance tuningHands on experience deploying, configuring and orchestrating software on Linux and Kubernetes, with proven ability to troubleshoot issues in both environmentsStrong experience with infrastructure as code (Ansible preferred) and best practicesProficient programming experience in PythonAbility to read, write and tune SQL queriesComfortable reading Java source code, tuning and debugging running JVMs.Familarity with data lakehouse technologies (e.g. Iceberg or Delta Lake) as well as query engine technologies (e.g. Dremio, Presto or Trino)Exposure to workflow orchestration tools like Airflow or DagsterA proactive mindset: you're not just fixing issues but preventing them.Comfortable working across teams, with minimal oversight. Our Tech Stack Data Tools: Hadoop (HDFS), Kafka, Dremio, Iceberg, Clickhouse, Spark, Airflow, FlinkInfrastructure Automation: Ansible, Puppet, Kubernetes (ArgoCD, Helm, Kustomize)Observability: Prometheus, Grafana, AlertManagerScripting: Python, Bash, SQLOthers: PCAP infrastructure About Us IMC is a research-driven trading firm where quantitative modeling, machine learning, and engineering shape how modern markets are traded. A stabilizing force in markets since 1989, we provide liquidity across trading venues, delivering the best outcome in value and risk management to investors. Using our own technology and capital, we build proprietary systems and algorithms that operate across global markets. Our researchers, traders, and engineers work as a collective, combining rapid experimentation, advanced infrastructure, and real-time feedback to turn insight into execution and execution into advantage.