Summary
✨ AI‑Generated
A technology team is seeking an experienced SRE to improve production reliability, implement monitoring solutions, automate operations, and support large-scale digital services.
Highlights
Work on reliability and scalability of high-traffic digital platforms with focus on automation, observability, and production excellence.
Description
Site Reliability Engineer - NO H1's / NO C2C
Location: Englewood, NJ
Duration: 12+ Months Contract
Work Arrangement: Onsite
Interview: In-person — Please Apply only IF you can do IN-PERSON INTERVIEW
Job Overview
We are looking for an experienced Site Reliability Engineer (SRE) to support and improve the reliability, performance, and observability of high-traffic digital platforms.
The ideal candidate will have strong experience with cloud infrastructure, CI/CD, monitoring, automation, and production support across web, mobile, and OTT environments.
Key Responsibilities
Support and enhance observability, including monitoring, logging, and alerting across production systems.Help define and maintain SLIs/SLOs for critical services.Evaluate services for production readiness.Work closely with development teams to identify reliability risks and improve system architecture.Automate operational processes, including CI/CD, incident response, and infrastructure provisioning.Participate in incident response and on-call rotations for critical services.Conduct post-incident analysis and drive continuous reliability improvements.Partner with security, infrastructure, and product teams to support performance, compliance, and operational excellence.Must-Have Skills
5+ years of experience managing and supporting high-traffic digital platforms.Strong experience with CI/CD pipelines and deployment automation.Hands-on experience with AWS and/or GCP.Strong scripting skills using Python, Bash, and/or Groovy.Experience with observability and monitoring tools such as:DatadogNew RelicAppDynamicsSimilar monitoring platformsUnderstanding of Web, Mobile, and OTT architectures.Experience supporting:Large-scale websitesMobile and OTT applicationsMicroservicesAPIsDistributed systemsExperience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or Chef.Familiarity with performance testing tools such as JMeter or k6.Hands-on experience with debugging tools such as Charles Proxy or Fiddler.Willingness to work onsite and participate in a 24/7 on-call rotation as needed.Preferred Qualifications
Experience with CDNs, such as Akamai.Experience with reverse proxies such as NGINX or Varnish.Exposure to video streaming platforms.Familiarity with application and infrastructure security controls and best practices.Certifications in SRE, DevOps, or Performance Engineering are a plus.