Senior Site Reliability Engineer

Bairesdev — Brazil · Posted ~2 days ago

Senior Full-time Remote

Skills

Site Reliability Engineering Kubernetes system architecture distributed systems production operations scalability reliability SRE cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

We are looking for a Senior Site Reliability Engineer to build and own reliable production services operating at massive scale. You will design and implement an agent safety layer around Kubernetes, reason about system architecture from small deployments to hundreds of millions of users, and operate services end-to-end in production. The role follows established SRE principles and emphasizes engineering ownership rather than simply administering infrastructure.

Highlights

Fully remote senior SRE role with end-to-end ownership of production services, large-scale architecture challenges, Kubernetes, and reliability engineering with substantial technical impact.

Description

At BairesDev®, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley. Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide. When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success. Senior Site Reliability Engineer (SRE) at BairesDev In this role, you'll build and own an agent safety layer that sits in front of Kubernetes, reasoning about system architecture at massive scale, from a handful of users up to hundreds of millions. Following the classic Google SRE model, you'll build and own a service end-to-end and run it reliably in production, not just administer infrastructure someone else designed. This is your opportunity to work on critical infrastructure where phased rollouts and canary releases are the standard, and where your architectural decisions directly protect production stability at scale. What You'll Do Build and own an agent safety layer sitting in front of Kubernetes, likely written in Go. Design systems that scale from a handful of users to hundreds of millions. Ensure changes reach production safely through phased rollouts and canary releases. Reason about distributed systems architecture and production reliability end-to-end. What We Are Looking For 5+ years of experience in Site Reliability Engineering. Strong experience with distributed systems design. Proficiency in Python or Golang. Experience with Kubernetes. Background in production reliability, including SLAs and SLOs. Experience with safe rollout practices such as canary or phased deployment. Advanced proficiency in English. Nice To Have Experience with observability or metrics tooling. How we do make your work (and your life) easier: 100% remote work (from anywhere). Excellent compensation in USD or your local currency if preferred Hardware and software setup for you to work from home. Flexible hours: create your own schedule. Paid parental leaves, vacations, and national holidays. Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent. Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities. Join a global team where your unique talents can truly thrive and make a significant impact! Apply now! #BD-PRIO-2026