ML Observability Engineer

Xpertdirect — Germany · Posted ~10 hours ago

Mid Full-time Hybrid

Skills

Python ML observability MLOps machine learning monitoring monitoring model performance Kubernetes OpenTelemetry Prometheus telemetry pipelines SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A growing AI-focused organization is seeking an ML Observability Engineer to build systems that provide deep visibility into production machine-learning workloads. You will develop monitoring and reliability tooling in Python, create telemetry pipelines with OpenTelemetry, build metrics and alerting with Prometheus, and integrate observability into Kubernetes-based ML infrastructure. The role sits at the intersection of MLOps, ML engineering, and SRE.

Highlights

Work at the intersection of MLOps, machine learning engineering, and site reliability engineering. The role offers hands-on ownership of production ML observability, including model performance, drift, latency, infrastructure behavior, telemetry, dashboards, and reliability tooling.

Description

ML Observability Engineer Berlin, Germany — Hybrid AI SaaS | ML Observability | MLOps | ML Reliability | AI Infrastructure Our client, a growing AI SaaS company based in Berlin, is looking for an ML Observability Engineer to build the systems that give engineering teams visibility into model performance, infrastructure behaviour, latency, drift, and production failures. You'll work at the intersection of MLOps, Site Reliability Engineering, and ML Engineering, helping teams understand not just whether their services are running—but whether their models are performing as expected. What You'll Work On • Build observability systems for production ML workloads • Monitor model performance, drift, latency, errors, and resource utilisation • Develop ML monitoring and reliability tooling in Python • Build telemetry pipelines using OpenTelemetry • Create metrics, dashboards, and alerting with Prometheus • Integrate observability across Kubernetes-based ML infrastructure • Track experiments, deployments, and model versions using MLflow • Define meaningful reliability indicators for production ML systems • Detect degradation and anomalous model behaviour before it impacts users • Improve incident investigation and root-cause analysis across models and infrastructure • Collaborate with ML, MLOps, and Platform Engineers to improve production reliability Core Skills • 4+ years in MLOps, ML Infrastructure, SRE, Platform Engineering, or similar roles • Python • MLflow • OpenTelemetry • Kubernetes • Prometheus • Model monitoring • Strong understanding of production ML systems and observability Nice to Have Grafana Evidently / Arize / WhyLabs or similar ML observability tooling Data and concept drift detection Distributed tracing SLI / SLO design PyTorch / TensorFlow Feature and data-quality monitoring AWS / GCP Terraform Incident management and postmortems LLM observability and evaluation