Senior / Principal Site Reliability Engineer

Oracle — United Kingdom · Posted ~3 hours ago

Lead Visa History ✓

Skills

Site Reliability Engineering Distributed systems Cloud infrastructure Production troubleshooting Automation Monitoring Telemetry Alerting Self-healing systems Reliability engineering Database Storage

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior-level SRE opportunity focused on improving the stability, scalability, performance, and operability of large-scale distributed cloud services. You will troubleshoot live production issues, build automation, develop telemetry and alerting, reduce operational risk, and create self-healing capabilities across multiple engineering environments.

Highlights

Hands-on senior engineering role focused on reliability, scalability, performance, and operability of critical distributed cloud infrastructure. The position emphasizes complex production problem-solving, automation, observability, and self-healing capabilities.

Description

Oracle Cloud Infrastructure (OCI) Core Infrastructure Engineering is looking for an experienced Principal Site Reliability Engineer (IC4) to help improve the reliability, performance, and operability of some of OCI’s most critical database and storage services. This is a hands-on engineering role for someone who enjoys solving complex production problems, building automation, and working with large-scale distributed cloud infrastructure. What you’ll be doing: • Improve the stability, performance, scalability, and reliability of database and storage services that form part of OCI’s core infrastructure • Troubleshoot complex production and live-service issues across highly distributed systems • Build automation and tooling to eliminate repetitive operational work and prevent recurring incidents • Develop telemetry, monitoring, alerting, and self-healing capabilities • Identify operational risks and improvements that can be applied across multiple services and engineering teams • Support release certification and help determine production readiness • Work closely with development teams on security, resiliency, capacity planning, performance, deployment, and release engineering • Participate in a 24×7 on-call rotation supporting critical OCI infrastructure What we’re looking for: • 8+ years of experience in Site Reliability Engineering and/or operating large-scale production systems • Strong troubleshooting experience across storage, networking, databases, and cloud services • Experience with major cloud platforms such as OCI, AWS, GCP, or Azure • Strong scripting and automation skills, particularly in Linux environments • Programming/infrastructure-as-code experience with technologies such as Python, Golang, and Terraform • Experience with Kubernetes, Docker, or similar container technologies • CI/CD experience with tools such as Jenkins, GitHub, Bitbucket, and Git • Experience with configuration management and monitoring/observability tools such as Grafana • Experience operating highly available 24×7 production environments • Strong communication, organizational, and problem-solving skills Important location & eligibility information: This position can be performed remotely within the UK. Candidates must already be based and residing in the UK and have valid authorization to work in the UK. Please also note that participation in an on-call rotation is an integral part of the role.