Senior Site Reliability Engineer

The Rework Group โ€” United States ยท Posted ~2 days ago

Senior

Skills

Site Reliability Engineering DevOps Incident response On-call operations Root cause analysis Observability AWS Terraform CloudFormation Infrastructure as Code CI/CD High-availability architecture Database optimization Data pipeline reliability EC2 Fargate Datadog Prometheus ELK

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Join a rapidly scaling technology organization as a senior reliability engineer and build resilient infrastructure for a growing user base. You will lead incident response, establish sustainable on-call practices, develop self-service observability, manage cloud infrastructure as code, improve deployment pipelines, and partner with product teams to design highly available systems. The role requires at least five years of SRE or DevOps experience, or seven years of software engineering experience with a strong infrastructure focus.

Highlights

Build reliability foundations from the ground up, own critical infrastructure challenges, establish scalable operational practices, and enable engineering teams to ship quickly with confidence.

Description

At The ReWork Group, we partner with high-growth startups and forward-thinking companies to build the future. To support our client's immediate predictive analytics endeavors, we're on the lookout for a Senior SRE. Our client sits at the intersection of two worlds that have long operated in parallel: the incumbent financial system, with its entrenched rails, opaque pricing, and gate kept access, and the decentralized economy, where programmable money, self-custody, and on-chain transparency are rewriting the rules of ownership and trust. They're scaling from thousands to millions users and they need the infrastructure to match. As a Senior SRE, you're not inheriting a mature system and maintaining the status quo. You're building the reliability foundation from the ground up, architecting resilient infrastructure, designing proactive monitoring that catches problems before users feel them, and standing up on-call processes that actually scale with the team. The challenges are real and immediate: database optimization, async workflow infrastructure, data pipeline reliability. You'll own them. And in doing so, you'll give the engineering team something invaluable โ€” the confidence to ship fast without breaking things. This is the kind of role where your fingerprints end up everywhere. What You'll Do: Lead incident response and establish sustainable on-call practices, including comprehensive runbooks, blameless postmortems, and systematic improvements that reduce MTTRDevelop and maintain self-service observability solutions using modern monitoring tools that provide actionable insights for troubleshooting and performance optimizationCreate and maintain infrastructure as code (using Terraform, CloudFormation) that allows for consistent, scalable, and secure cloud environments on AWSPartner closely with feature teams to architect resilient infrastructure for critical components (databases, networking, async workflows, data pipelines) that scale seamlesslyWork closely with DevX to design and implement robust CI/CD pipelines with advanced deployment strategies (blue/green, canary) that enable teams to ship confidently and rapidlyAdvocate for best practices early in feature design, ensuring we design with reliability in mind and future-proof our services What You'll Bring: 5+ years in SRE or DevOps โ€” or 7+ years in software engineering with a serious infrastructure focus and the scars to prove itYou've led incident response for high-availability production systems โ€” you run tight RCAs, you drive blameless postmortems, and you leave every incident with a team that's smarter than beforeYou've designed highly available deployment architectures across multiple targets โ€” EC2, Fargate, and beyond โ€” with real expertise in auto-scaling, health checks, and graceful degradation when things get hardYou've implemented monitoring and observability solutions that actually get used โ€” Datadog, Prometheus, ELK, or comparable โ€” and you've made the case internally for why observability isn't optionalDeep AWS fluency and a strong infrastructure-as-code practice โ€” Terraform is your default, not your fallbackYou've built and improved CI/CD pipelines that give engineering teams the confidence to ship fast and reliably