Observability & Site Reliability Engineer

Ubique Systems — Germany · Posted ~20 hours ago

Senior Contract Remote

Skills

Site Reliability Engineering Observability Prometheus Grafana Loki OpenTelemetry Jaeger PromQL LogQL AlertManager Cloud environments Monitoring Distributed tracing

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced SRE and observability role responsible for building and operating monitoring and telemetry platforms across distributed cloud environments. You will work with metrics, logs, tracing, dashboards, alerts, and reliability improvements using a modern open-source observability stack.

Highlights

Fully remote role with a competitive salary, focused on modern observability technologies and the opportunity to build central monitoring and telemetry platforms for distributed cloud environments.

Description

Observability & SRE Engineer Location: Remote – Germany preferred Contract: 1-Year Fixed Term Contract Salary: Up to €80,000 – Flexible for strong candidates About the Role We are looking for an experienced Observability & SRE Engineer to build and operate central monitoring and telemetry platforms for distributed cloud environments. You will work with metrics, logs, distributed tracing and alerting, helping engineering teams understand system health, troubleshoot issues and improve reliability. The role involves working with Prometheus, Grafana, Loki, OpenTelemetry and Jaeger, including environments where operations need to work independently or in air-gapped environments. What You'll Do Build and maintain observability and monitoring platforms using Prometheus, Grafana, Loki and Jaeger.Create monitoring dashboards and alerts using PromQL and LogQL.Configure AlertManager for alert routing and notification.Implement OpenTelemetry Collectors and SDKs for application monitoring and distributed tracing.Support OTLP telemetry collection and trace propagation.Implement distributed tracing using Jaeger.Configure secure log forwarding and integration with external SIEM platforms.Work with security event logging using Syslog / CEF.Support monitoring and telemetry solutions for distributed and air-gapped environments.Automate deployments and configuration using Infrastructure-as-Code and scripting.Troubleshoot monitoring, logging, alerting and performance issues.Work closely with SRE, DevOps, Cloud, Security and Platform teams. Essential Skills Strong hands-on experience with Prometheus.Strong experience with Grafana and dashboard development.Experience with PromQL.Experience with Loki / LogQL or other centralised log aggregation platforms.Strong understanding of OpenTelemetry, including Collectors, SDKs and OTLP.Experience with Jaeger / distributed tracing.Experience with AlertManager and alert routing.Understanding of SIEM integration and security event logging.Experience with Syslog, CEF or similar security logging formats.Strong troubleshooting and monitoring skills.Scripting experience with Go, Python, Shell and/or YAML.Clearance Requirement Candidates must be able to obtain/clear: Public Sector ClearanceNdK – Nachweis der Kundigkeit If you're an experienced SRE, Observability Engineer, DevOps Engineer or Platform Engineer with strong Prometheus/Grafana experience, we'd be interested in hearing from you. Please apply or send your latest CV via LinkedIn message. #SRE #SiteReliabilityEngineer #Observability #ObservabilityEngineer #DevOps #PlatformEngineer #Prometheus #Grafana #OpenTelemetry #Jaeger #Loki #Kubernetes #CloudNative #Monitoring #DistributedTracing #SIEM #DevOpsJobs #SREJobs #GermanyJobs #RemoteJobs #Germany