Summary
✨ AI‑Generated
A Senior Site Reliability Engineer is sought to bridge software engineering and systems operations, building and protecting highly available production environments. You will automate infrastructure, optimize deployment pipelines, implement observability, establish reliability objectives, and improve the safety and speed of software delivery.
Highlights
Own production reliability by automating operations, scaling infrastructure, and improving availability, resilience, and performance. The role offers hands-on work with modern cloud infrastructure, Infrastructure as Code, CI/CD, observability, and reliability engineering practices.
Description
We are seeking a skilled and proactive Senior Site Reliability Engineer (SRE) to join our engineering team.
In this role, you will bridge the gap between software development and systems operations, applying software engineering principles to automate operations, scale infrastructure, and ensure systems are highly available, resilient, and performant.
Your mission is to build, run, and protect the production environments that power our applications, minimizing downtime and enabling rapid, safe software deployment.
Responsibilities
Design, build, and maintain cloud infrastructure using modern Infrastructure as Code practices such as Terraform and CloudFormationBuild and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasksDesign and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and DatadogEstablish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)Respond to production incidents and lead troubleshooting efforts to restore servicesConduct blameless post-mortems to identify root causes and prevent recurrencePartner with software developers to optimize system performance and plan capacityEnsure services can scale to handle growth and traffic spikes
Requirements
3+ years of experience in systems administration, DevOps, or systems-focused software developmentProficiency in at least one scripting or programming language such as Python, Bash, Go, or RustExperience with public cloud providers such as AWS, Azure, or GCP, along with containerization tools such as Docker and KubernetesUnderstanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLSPassion for automation, eliminating toil, and building resilient systems that fail gracefullyAdvanced proficiency in English (C1+)
We offer
International projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer.
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.