Summary
✨ AI‑Generated
Join a senior SRE role focused on making an AI-accelerated engineering organization faster and more reliable. You will evolve cloud infrastructure, automate validation, strengthen release safety, apply SLOs and observability, improve incident response, and reduce manual operational work while supporting production systems at scale.
Highlights
Shape reliability practices for an AI-accelerated engineering team, using automation, observability, safe deployment strategies, and rapid feedback to improve delivery speed while reducing operational risk and manual work.
Description
The Role
Help Rabbit build and ship faster with AI — safely, securely and reliably.
Rabbit runs automated cost optimization across enterprise Google Cloud environments.
When we change a customer's BigQuery reservations or rightsize their GKE clusters, those changes need to be correct and dependable.
Reliability is central to the trust customers place in our product.
Our foundation is already in place: logging, alerting, automated deployment and Terraform-managed infrastructure.
Your mission is to evolve that foundation for an AI-accelerated engineering team: turn faster implementation into faster, dependable delivery through automated validation, safe releases and rapid feedback.
You'll apply proven SRE practices — SLOs, observability, incident response and deployment safety — to AI-assisted development and agent-driven workflows.
The goal is to increase how quickly the team can deliver verified improvements, while controlling production risk and reducing manual operational work.
What You'll Do
Make AI-assisted delivery faster and safer.
Build automated validation, progressive rollout and recovery mechanisms that let engineers and agents move quickly with clear checks before and after changes reach production.Make reliability measurable.
Define and operationalize SLOs, SLIs and error budgets, and use them to guide practical decisions about delivery speed, stability and reliability work.Automate workflows with AI agents.
Identify repetitive operational work and build reusable agent-driven workflows for alert triage, incident investigation, routine maintenance and reporting.
Add verification and human approval where needed, and measure the reduction in manual effort.Improve observability and feedback.
Evolve logging, metrics, tracing and alerting so failures are detected early and changes can be traced, investigated and verified.Turn incidents into lasting improvements.
Improve runbooks, investigation and blameless postmortems, and translate recurring problems into tests, safeguards and automation.Keep infrastructure reproducible.
Extend our Terraform and delivery tooling so environments remain consistent and changes stay reviewable as the platform grows.Improve GCP reliability and efficiency.
Strengthen our cloud infrastructure, networking, access controls and capacity management, balancing performance, reliability and cost.Use AI to accelerate reliability engineering itself.
Build maintainable tooling in Go, Python or a comparable language, and use agents to accelerate investigation, implementation, testing and documentation while verifying their outputs.
How We Work — AI-First, Agentic by Default
AI-assisted engineering is an expectation of this role, not an optional experiment.
Claude Code, Cursor and agent-driven workflows are part of how we work, including infrastructure and reliability engineering.
We want someone who actively looks for ways to increase engineering speed with AI and makes those improvements safe to repeat.
That means shorter feedback loops, automated checks, traceable changes and recovery paths — not simply generating more code.
You remain accountable for engineering judgment: what to automate, how to verify it, when human approval is needed and when a change should be stopped or rolled back.
Success means faster delivery of reliable improvements, less repetitive work and a platform the team can trust.
What You'll Bring
Must-have
6+ years in SRE, production engineering or infrastructure-heavy backend roles, with hands-on ownership of production systems.Strong production GCP experience, including Cloud Run, networking and IAM.
Hands-on Google Cloud experience is required and will be assessed during the interview process.Infrastructure-as-code fluency with Terraform, plus solid experience in CI/CD and deployment safety.Strong observability and troubleshooting skills: you can make systems debuggable, identify root causes and verify that a fix works.Coding ability in Go, Python or a comparable language, with experience building maintainable operational tooling.Experience leading production incident investigation and driving follow-up improvements that prevent recurrence.Practical fluency with AI coding agents and the ability to critically review, test and validate their work.
You are motivated to make AI-assisted engineering faster and more dependable.Strong written English and the discipline to collaborate asynchronously with a distributed team.
Nice to have
Kubernetes / GKE experience, including deploying, operating, debugging and scaling containerized services.Hands-on Datadog experience, including dashboards, monitors, logs, APM and distributed tracing.Experience improving delivery speed and safety through progressive delivery, policy-as-code and automated rollback.GCP cost management or FinOps experience.Security experience is a plus.
Why This Role
Shape how an AI-first engineering team scales.
Build the reliability practices and automation that let Rabbit turn faster development into dependable customer outcomes.Work on systems with real customer impact.
Rabbit operates inside enterprise GCP environments, where reliability and security directly affect customer trust.Own meaningful improvements.
Work in a small team with short decision paths and end-to-end ownership, supported by appropriate review and production safeguards.Use AI as an engineering multiplier.
Apply agentic tools to infrastructure, delivery and operations, and help define how we measure and improve their impact.