Site Reliability Engineer

Epam Systems โ€” Chile ยท Posted ~1 day ago

Mid Visa History โœ“

Skills

Infrastructure as Code CI/CD cloud infrastructure Docker Kubernetes Linux/Unix administration networking incident response monitoring automation Terraform CloudFormation Prometheus Grafana Datadog Python Bash Go Rust AWS Azure GCP Linux TCP/IP DNS HTTP SSL/TLS

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Join an engineering team responsible for building resilient, scalable, and highly available production platforms. You will automate infrastructure and deployments, develop observability systems, define reliability objectives, lead incident response, and collaborate with developers on performance and capacity. The role requires at least two years of relevant experience, cloud and container expertise, strong Linux and networking knowledge, and advanced English.

Highlights

Work on international engineering projects with global teams, comprehensive healthcare and financial benefits, paid leave, extensive training and certification resources, and opportunities for community involvement.

Description

Our engineering team is looking to add a skilled and proactive Site Reliability Engineer (SRE). This position serves as the connective link between software development and systems operations. Software engineering principles will be applied to automate operations, scale infrastructure, and keep systems highly available, resilient, and performant. The core mission involves building, running, and safeguarding the production environments that power our applications, keeping downtime to a minimum while enabling fast, safe software deployment. Responsibilities Design, build, and maintain cloud infrastructure using modern Infrastructure as Code practices such as Terraform and CloudFormationBuild and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasksDesign and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and DatadogEstablish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)Respond to production incidents and lead troubleshooting efforts to restore servicesConduct blameless post-mortems to identify root causes and prevent recurrencePartner with software developers to optimize system performance and plan capacityEnsure services can scale to handle growth and traffic spikes Requirements 2+ years of experience in systems administration, DevOps, or systems-focused software developmentProficiency in at least one scripting or programming language such as Python, Bash, Go, or RustExperience with public cloud providers such as AWS, Azure, or GCP, along with containerization tools such as Docker and KubernetesUnderstanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLSPassion for automation, eliminating toil, and building resilient systems that fail gracefullyAdvanced proficiency in English (C1+) We offer International projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.