Senior Site Reliability Engineer

Hostpapa — Bulgaria · Posted ~3 hours ago

Senior Full-time

Skills

SRE SLI/SLO error budgets incident response Terraform Kubernetes Azure AWS Python Go cloud infrastructure observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering opportunity focused on building and operating highly available cloud-native platforms. The role involves improving scalability, automation, monitoring, incident response processes, and infrastructure reliability using modern cloud and container technologies.

Highlights

Senior hands-on reliability role working on large-scale cloud platforms with ownership of automation, scalability, observability, and production reliability improvements.

Description

We’re looking for a Senior SRE to drive reliability for CloudBlue’s multi-tenant, multi-cloud SaaS platform. What you bring: Proven experience owning reliability, scalability, and observability for enterprise-grade, multi-tenant SaaS platforms at scale, including defining SLIs/SLOs and managing error budgets Experience acting as a senior incident responder; leading incident coordination, driving blameless postmortems, and reducing MTTR and incident frequency Hands-on Terraform experience authoring reusable modules (not just consuming them), ideally on Azure Strong Kubernetes day-2 operations experience (upgrades, networking, RBAC, and hands-on troubleshooting beyond dashboards)Deep infra-level expertise on a major cloud platform (Azure or AWS) Demonstrated ability to build end-to-end automation in Python and/or Go to reduce operational toil at scale This is a senior, hands-on role working across Kubernetes-based platforms and global production systems. ** This role is open to candidates based in the European Union, in line with current operational and business requirements**