Summary
✨ AI‑Generated
Join a technology company to support cloud services, build automation, improve monitoring, and incident response. Role involves operational tooling and supporting production services at scale. Collaborate with engineering teams to enhance reliability and service resilience.
Highlights
Opportunity to work on OCI storage and data services, build automation, improve monitoring and incident response. Collaborate with engineering teams to enhance reliability and operational efficiency.
Description
We’re hiring an SRE to support new OCI storage and data services, with a focus on reliability, operations, automation, and scaling production services.
This role is ideal for someone with a strong SRE or infrastructure engineering background who enjoys building operational tooling, improving service reliability, and supporting critical cloud services at scale.
Key responsibilities:
Support OCI storage and data-oriented services in productionProvide operational coverage across UK and non-US time zonesBuild automation to reduce manual operational workImprove monitoring, dashboards, telemetry, and alertingSupport incident response, troubleshooting, and root cause analysisBuild tools and processes that improve reliability and operational efficiencyPartner closely with engineering teams to improve service readiness and resilience
Ideal experience:
Strong SRE, infrastructure, or production engineering backgroundautomation skillsExperience supporting large-scale production servicesHands-on experience with observability, dashboards, alerting, and incident managementAbility to diagnose complex issues and drive long-term reliability improvementsBackground in storage, databases, streaming, distributed systems, or data-oriented services
Relevant technical backgrounds may include:
Database internals, such as MySQL or OracleBlock storage, file systems, SAN, or distributed storageStreaming, replication, or key-value systemsLarge-scale platforms with similarities to Kafka, DynamoDB, or database-like systems
This is a great opportunity for someone who wants to work on high-impact cloud infrastructure, improve the reliability of critical services, and help mature systems as they scale.