Kubernetes Reliability Engineer

Xpertdirect — Austria · Posted ~4 hours ago

Hybrid

Skills

Kubernetes Site Reliability Engineering Platform engineering Go Prometheus OpenTelemetry Terraform Linux Observability SLIs/SLOs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work at the intersection of Kubernetes, SRE, and platform engineering to make large-scale cloud environments more resilient and operable. You will build Go-based tooling, establish reliability standards, develop observability with modern telemetry technologies, automate infrastructure with Terraform, and troubleshoot complex production issues.

Highlights

Hybrid role focused on large-scale cloud infrastructure reliability, with hands-on work in Kubernetes, automation, observability, incident detection, and platform tooling.

Description

Kubernetes Reliability Engineer Vienna, Austria — Hybrid Cloud Infrastructure | Kubernetes | Site Reliability Engineering | Platform Engineering Our client, a growing Cloud Infrastructure company based in Vienna, is looking for a Kubernetes Reliability Engineer to improve the resilience, observability, and operability of large-scale Kubernetes environments. You'll work at the intersection of Kubernetes Engineering, SRE, and Platform Engineering, building the automation and observability systems that keep critical cloud infrastructure reliable in production. What You'll Work On • Improve reliability and resilience across production Kubernetes clusters • Build platform and operational tooling in Go • Define and implement SLIs, SLOs, and reliability standards • Develop observability using Prometheus and OpenTelemetry • Automate infrastructure provisioning and configuration with Terraform • Investigate complex Kubernetes and Linux production issues • Improve alerting, incident detection, and operational visibility • Automate repetitive operational and recovery workflows • Analyse capacity, performance, and infrastructure bottlenecks • Work with engineering teams to improve application reliability on Kubernetes Core Skills • 4+ years in SRE, Platform Engineering, Kubernetes Engineering, Cloud Infrastructure, or similar roles • Strong Kubernetes experience in production environments • Go • Prometheus • OpenTelemetry • Terraform • Linux • Strong understanding of distributed systems and production reliability Nice to Have Grafana Kubernetes Operators / Controllers Helm Argo CD / GitOps AWS / GCP / Azure eBPF Service meshes Incident management and postmortems Capacity planning and performance engineering Experience operating multi-cluster or large-scale Kubernetes environments