Summary
✨ AI‑Generated
A growing AI-focused organization is seeking an ML Observability Engineer to build systems that provide deep visibility into production machine-learning workloads. You will develop monitoring and reliability tooling in Python, create telemetry pipelines with OpenTelemetry, build metrics and alerting with Prometheus, and integrate observability into Kubernetes-based ML infrastructure. The role sits at the intersection of MLOps, ML engineering, and SRE.
Highlights
Work at the intersection of MLOps, machine learning engineering, and site reliability engineering. The role offers hands-on ownership of production ML observability, including model performance, drift, latency, infrastructure behavior, telemetry, dashboards, and reliability tooling.
Description
ML Observability Engineer
Berlin, Germany — Hybrid
AI SaaS | ML Observability | MLOps | ML Reliability | AI Infrastructure
Our client, a growing AI SaaS company based in Berlin, is looking for an ML Observability Engineer to build the systems that give engineering teams visibility into model performance, infrastructure behaviour, latency, drift, and production failures.
You'll work at the intersection of MLOps, Site Reliability Engineering, and ML Engineering, helping teams understand not just whether their services are running—but whether their models are performing as expected.
What You'll Work On
• Build observability systems for production ML workloads
• Monitor model performance, drift, latency, errors, and resource utilisation
• Develop ML monitoring and reliability tooling in Python
• Build telemetry pipelines using OpenTelemetry
• Create metrics, dashboards, and alerting with Prometheus
• Integrate observability across Kubernetes-based ML infrastructure
• Track experiments, deployments, and model versions using MLflow
• Define meaningful reliability indicators for production ML systems
• Detect degradation and anomalous model behaviour before it impacts users
• Improve incident investigation and root-cause analysis across models and infrastructure
• Collaborate with ML, MLOps, and Platform Engineers to improve production reliability
Core Skills
• 4+ years in MLOps, ML Infrastructure, SRE, Platform Engineering, or similar roles
• Python
• MLflow
• OpenTelemetry
• Kubernetes
• Prometheus
• Model monitoring
• Strong understanding of production ML systems and observability
Nice to Have
Grafana
Evidently / Arize / WhyLabs or similar ML observability tooling
Data and concept drift detection
Distributed tracing
SLI / SLO design
PyTorch / TensorFlow
Feature and data-quality monitoring
AWS / GCP
Terraform
Incident management and postmortems
LLM observability and evaluation