Summary
✨ AI‑Generated
A technology team is seeking a senior SRE to maintain highly available platforms, improve system reliability, lead incident response, and build resilient cloud-based services.
Highlights
Remote senior role focused on reliability engineering, automation, observability, and operating large-scale production systems.
Description
Senior Site Reliability Engineer (SRE)
REMOTE
Full-Time | Permanent
Salary: $125,000 – $145,000 USD annually
The Opportunity
We are looking for an experienced Senior Site Reliability Engineer (SRE) to join a growing technology team responsible for highly available, mission-critical platforms operating 24/7.
You’ll work on modern, cloud-native systems that process high volumes of real-time transactions and data.
This is a hands-on senior position where you'll take ownership of system reliability, lead incident response, improve observability and automation, and work closely with engineering and infrastructure teams to build resilient, scalable services.
What You’ll Be Doing
Monitor and maintain the availability, reliability and performance of business-critical production systems.Define and track SLOs, SLIs, error budgets and reliability metrics.Lead production incident response, coordinating technical teams through diagnosis and resolution.Conduct root cause analysis and lead blameless post-incident reviews, ensuring follow-up actions are completed.Build automation using Python, Bash or Go to reduce manual operational work.Develop and improve monitoring, logging, tracing and alerting using technologies such as Prometheus, Grafana, OpenTelemetry, ELK and Datadog.Identify performance bottlenecks and troubleshoot issues across both infrastructure and application layers.Support capacity planning and infrastructure scaling for growing transaction volumes.Work with stateful and distributed technologies including SQL databases and event-streaming platforms such as Kafka or NATS.Collaborate with software engineering, infrastructure and operations teams to improve overall platform reliability.Help design, implement and test disaster recovery and business continuity processes.Improve operational documentation, processes and SRE best practices.Mentor engineers and help establish a strong reliability engineering culture.What We’re Looking For
5+ years of experience within Site Reliability Engineering, DevOps, Platform Engineering or a similar production-focused role.Experience supporting mission-critical, highly available or large-scale production environments.Strong experience with at least one major cloud platform: AWS, Azure or GCP.Hands-on experience with Kubernetes and Docker.Strong understanding of observability and monitoring using tools such as Prometheus, Grafana, ELK, Datadog or OpenTelemetry.Strong scripting skills using Python, Bash and/or Go.Solid Linux administration experience, with exposure to Windows environments beneficial.Experience leading production incidents, root cause analysis and post-incident reviews.Understanding of SLOs, SLIs, error budgets and reliability engineering principles.Familiarity with relational databases such as PostgreSQL or MySQL and broader distributed-system concepts.Strong troubleshooting and problem-solving ability.Excellent communication skills and the ability to collaborate across engineering and operational teams.Nice to Have
Terraform, Ansible and/or Helm.GitOps experience using tools such as Argo CD.CI/CD and pipeline-as-code experience, including GitHub Actions or Dagger.Experience with Kafka, NATS or other event-streaming technologies.Experience supporting high-volume transactional or payment systems.Exposure to Intelligent Transportation Systems (ITS), tolling, traffic management, sensor processing or other real-time infrastructure.Experience designing and testing disaster recovery strategies.Why Join?
This is an opportunity to work on complex, real-world technology where reliability directly matters.
You’ll have significant ownership across production reliability, observability, automation and incident management while helping shape SRE practices as the platform continues to scale.
Salary: $125,000 – $145,000 USD annually, dependent on experience.