Principal Site Reliability Engineer

Entain — Austria · Posted ~2 hours ago

Lead Visa History ✓

Skills

Site Reliability Engineering Cloud infrastructure Hybrid infrastructure Observability Monitoring Logging Alerting Cloud SRE DevOps Infrastructure as Code

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior engineering role responsible for defining reliability architecture, improving system performance, establishing operational best practices, and mentoring technical teams without direct people management.

Highlights

Highly technical leadership role shaping reliability standards, cloud architecture, scalability practices, and engineering excellence across an organization.

Description

Job DescriptionWe are seeking a Principal Site Reliability Engineer(SRE) to set the technical direction for the reliability, performance, and scalability of our hybrid infrastructure. As the most senior technical authority in the SRE function, you will define architecture and engineering standards across cloud and on-premise environments, drive the adoption of modern SRE practices, and ensure operational excellence. This is a hands-on, deeply technical role (~90%) with significant cross-organizational influence and technical mentorship responsibilities (~10%). You will operate as a force multiplier, raising the technical bar for the entire engineering organization rather than managing a team directly. Technical Responsibilities (~90%) Architect scalable, secure, and cost-efficient cloud and hybrid solutions, and establish the patterns and reference implementations other teams build on.Define the observability strategy across the organization: monitoring, logging, alerting, SLIs/SLOs and error budgets using Datadog, Elastic, OpenTelemetry, and similar tooling.Set standards for high availability, disaster recovery, and backup across hybrid environments, and validate them through resilience and failure testing.Partner with development, security, and platform teams to shape deployment pipelines (CI/CD) and GitOps workflows at scale.Establish and maintain organization-wide technical documentation, runbooks, architectural decision records, and operational standards.Identify systemic reliability bottlenecks, lead complex root-cause analysis, and drive down operational load through automation and platform improvements.Troubleshoot and resolve the most complex, high-impact production issues, and lead major incident response.Provide deep 3rd line technical support and participate in the on-call rotation as an escalation point.Technical Leadership & Influence (~10%) Act as a technical mentor and role model for SREs and engineers across teams, raising the overall engineering bar through coaching and knowledge-sharing.Lead structured up skilling and enablement across Windows, Linux, AWS, and Kubernetes.Influence and align Product, Engineering, Security, Platform, and Operations teams around reliability goals and technical direction.Shape the SRE roadmap together with SRE Leadership and SRE Guild and drive prioritisation of key reliability and platform initiatives.Communicate technical strategy and trade-offs clearly to both engineering teams and senior stakeholders, and drive accountability for reliability outcomes.Champion a culture of operational excellence, ownership, and continuous improvement across the organization.