Summary
✨ AI‑Generated
A Site Reliability Engineer is sought to design, automate, and operate highly available production infrastructure at scale. The role focuses on Kubernetes, AWS, Terraform, cloud networking, observability, incident response, distributed systems, and automated delivery workflows, with ownership of service-level objectives and continuous reliability improvements. The position is remote.
Highlights
Remote SRE role with direct ownership of reliability, production readiness, availability, and performance. Provides hands-on work with AWS, Kubernetes, cloud networking, observability, infrastructure as code, and modern deployment automation.
Description
About The Role
The Site Reliability Engineer will design, automate, and operate the infrastructure supporting production services at scale.
The role focuses on Kubernetes-based platforms, cloud networking, observability, incident response, and the reliability of distributed systems running across AWS environments.
Working with application engineers, security, and platform teams, the engineer will turn operational requirements into resilient systems and repeatable delivery workflows.
The role is remote from New York, NY, with direct ownership of service-level objectives, production readiness, and continuous improvements to availability and performance.
Key Responsibilities
Design and operate highly available production infrastructure across AWS using Kubernetes, Terraform, and infrastructure-as-code best practicesBuild and maintain CI/CD pipelines with tools such as GitHub Actions, Argo CD, or Jenkins to automate testing, deployments, rollbacks, and environment provisioningDefine and enforce service-level objectives, error budgets, and production readiness standards for critical servicesImplement observability using Prometheus, Grafana, OpenTelemetry, and centralized logging platforms to reduce detection and resolution timeLead incident response, coordinate technical remediation, and produce blameless postmortems with measurable follow-up actionsImprove platform resilience through capacity planning, performance testing, disaster recovery exercises, and automated failure remediationPartner with development teams to improve application architecture, containerization, security controls, and operational runbooks
What We Are Looking For
3–8 years of experience in site reliability engineering, DevOps, platform engineering, or a related production infrastructure roleStrong hands-on experience with AWS services such as EC2, EKS, IAM, VPC, S3, RDS, and CloudWatchProficiency with Kubernetes, Docker, Terraform, and Linux systems administration in production environmentsExperience building and operating CI/CD pipelines and managing deployment strategies such as canary releases, blue-green deployments, and automated rollbacksWorking knowledge of observability, distributed systems, networking, and incident management practices, including SLOs, SLIs, and error budgetsProficiency in Python, Go, or Bash for automation, tooling, and operational workflows; bachelor’s degree in computer science, engineering, or equivalent practical experienceBonus: Experience with Argo CD, Helm, Prometheus, Grafana, OpenTelemetry, service meshes, multi-region architectures, or chaos engineering