Site Reliability Engineer

Epam Systems — Colombia · Posted ~1 week ago

Full-time Visa History ✓

Skills

Site Reliability Engineering Cloud infrastructure Infrastructure as Code Automation Observability Incident response Production operations Scalability Reliability engineering Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a global engineering environment as an SRE focused on making production systems reliable, scalable, and safe. You will design cloud infrastructure with Infrastructure as Code, automate operational workflows, improve observability, respond to incidents, and help teams deliver software quickly and safely.

Highlights

Work on production reliability, scalability, and safety while bridging software development and operations. The role emphasizes automation, observability, infrastructure as code, incident response, and enabling fast, safe releases in an international engineering environment.

Description

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential. We are seeking a proactive Site Reliability Engineer to strengthen the reliability, scalability, and safety of production environments. You will bridge software development and operations through automation, observability, and incident response—apply now to help reduce downtime and enable fast, safe releases. Responsibilities Design and maintain cloud infrastructure using Infrastructure as Code practicesBuild and optimize CI/CD pipelines to automate deployments and operational workflowsImplement logging, monitoring, and alerting to improve observability and reliabilityDefine and track Service Level Objectives and Service Level Indicators with clear reportingRespond to production incidents and drive rapid service restorationLead blameless post-mortems to identify root causes and prevent recurrencePartner with engineers to improve performance, scalability, and capacity planningAutomate repetitive operational tasks to reduce toil and operational riskHarden production environments to improve resilience and safe change practices Requirements 2+ years of experience in site reliability engineering, DevOps, or systems administrationHands-on experience with Infrastructure as Code using Terraform or CloudFormationHands-on experience building and improving CI/CD pipelines for automated deploymentsStrong troubleshooting and incident response leadership skills in production environmentsSolid project skills to coordinate reliability work with software development teamsProficiency in scripting or programming with Python, Bash, Go, or RustCloud platform experience with AWS, Azure, or GCPContainerization experience with Docker and KubernetesDeep Linux/Unix administration knowledge and networking fundamentals (TCP/IP, DNS, HTTP, SSL/TLS)Strong communication skills with a reliability mindset focused on automation and reducing toilAdvanced English proficiency (C1, Advanced) Nice to have Experience with Prometheus, Grafana, or Datadog We offer International projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.