Summary
✨ AI‑Generated
Join an engineering team responsible for highly available, scalable systems. You’ll work across cloud infrastructure, automation, observability, and deployment to improve reliability, reduce downtime, and operate resilient distributed systems at scale.
Highlights
Build highly available and scalable systems while solving complex infrastructure challenges, automating operations, improving performance, and strengthening reliability.
Description
Site Reliability Engineering (SRE)
This role is for one of our clients
Location: New York City, NY
Salary Range: $150,000 – $250,000 (Base)
About the Role
We are seeking an experienced Site Reliability Engineer (SRE) to join our engineering team in New York City.
In this role, you will be responsible for building and maintaining highly available, scalable, and reliable systems.
You will work across infrastructure, cloud platforms, automation, observability, and deployment processes to ensure our applications operate efficiently and reliably at scale.
You will collaborate closely with software engineers, DevOps teams, and infrastructure stakeholders to improve system performance, automate operational processes, strengthen reliability, and minimize downtime.
This role is ideal for engineers who enjoy solving complex infrastructure challenges, automating repetitive processes, and building resilient distributed systems.
Requirements
Required Qualifications
Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent experience)
3+ years of professional experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related field
Strong experience with cloud platforms such as AWS, Azure, or Google Cloud Platform
Hands-on experience with Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools
Experience designing and maintaining CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar technologies
Strong knowledge of Docker, Kubernetes, and container orchestration
Experience administering Linux-based systems and writing automation scripts using Bash, Python, or PowerShell
Experience with monitoring, observability, and logging tools such as Prometheus, Grafana, ELK, Datadog, or similar platforms
Strong understanding of system reliability, availability, scalability, and performance
Experience troubleshooting production systems and resolving complex infrastructure issues
Strong automation and problem-solving skills
Ability to work collaboratively with software engineering and infrastructure teams
Preferred Qualifications
Hands-on experience operating Kubernetes in production environments
Experience with Site Reliability Engineering practices, including SLIs, SLOs, SLAs, and error budgets
Knowledge of networking, cloud security, IAM, and infrastructure best practices
Experience implementing GitOps workflows using tools such as Argo CD or Flux
Familiarity with microservices and distributed systems architecture
Experience with incident response, root cause analysis, and production troubleshooting
Experience building scalable and highly available systems in cloud environments
Familiarity with capacity planning, disaster recovery, and business continuity strategies
Experience with observability, alerting, and performance monitoring at scale
Compensation & Benefits
Base Salary: $150,000 – $250,000, depending on experience and qualifications
Competitive equity or bonus opportunities
Comprehensive health, dental, and vision benefits
Generous paid time off and holidays
Professional development and learning opportunities
Opportunity to work with modern cloud and infrastructure technologies
Collaborative and innovative engineering environment in New York City