Site Reliability Engineer

Avance Services — Germany · Posted ~22 hours ago

Senior Full-time

Skills

Prometheus PromQL Grafana Loki Jaeger OpenTelemetry OTLP AlertManager LogQL SIEM integration Syslog CEF air-gapped infrastructure cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an SRE to build and operate central observability platforms across distributed cloud infrastructure. You will implement monitoring, logging, distributed tracing, alerting, and secure SIEM integrations, with a particular focus on autonomous air-gapped environments.

Highlights

Build and operate modern observability infrastructure for distributed cloud environments, including metrics, logs, tracing, and alerting. The role offers technically demanding work involving secure, autonomous, air-gapped environments and public-sector infrastructure.

Description

Position Overview Security Clearance & Vetting Level: Public Sector Clearance + NdK (Nachweis der Kundigkeit)The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. This role implements metric collection, log aggregation, distributed tracing, and standalone alerting tailored for Air-Gap operations. Key Responsibilities Monitoring Stack Setup: Deploy and maintain Prometheus, Grafana, Loki, and Jaeger stacks as code.Dashboards & Alerting: Develop PromQL/LogQL monitoring dashboards; configure AlertManager routing and inhibition for autonomous Air-Gap operations.Distributed Tracing: Implement OpenTelemetry Collectors and SDKs for auto-instrumentation and OTLP trace propagation.SIEM Integration: Configure secure log export and security event management (CEF/Syslog) for external SIEM platforms.Technical Qualifications & Skills Must Have:Deep expertise in Prometheus (PromQL, ServiceMonitor, Federation, Remote Write) and Grafana.Hands-on experience with Loki log aggregation and AlertManager routing.Proficiency with OpenTelemetry (Collectors, SDKs, OTLP) and Jaeger distributed tracing.Knowledge of SIEM integrations and security event logging.Scripting skills in Go, Python, Shell, and YAML.