Summary
✨ AI‑Generated
A senior Site Reliability Engineering role supporting large-scale enterprise production systems in a highly available distributed environment. You’ll respond to critical incidents, troubleshoot complex systems, lead major incident bridges when needed, improve observability and automation, and drive continuous reliability improvements. The role follows an evening-to-overnight Eastern Time schedule and is remote within Latin America.
Highlights
Fully remote senior SRE work supporting large-scale, highly available enterprise systems, with hands-on ownership of reliability, automation, observability and major incident response.
Description
Senior Site Reliability Engineer (Sr.
SRE)
Location: LATAM – Remote
Work Hours: 5:00 PM – 1:00 AM EST,
Experience: 8+ years
About the Role
We are seeking a Senior Site Reliability Engineer (SRE) to support large-scale enterprise production applications in a highly available and distributed environment.
This is a hands-on role focused on production incident response, troubleshooting, reliability engineering, automation, observability, and continuous improvement.
The ideal candidate will have strong technical troubleshooting skills, excellent communication abilities, and experience working in high-pressure production environments.
Key Responsibilities
Respond to production alerts and incidents, assess severity, and initiate appropriate mitigation actions.Troubleshoot complex distributed systems and identify root causes of production issues.Participate in and lead major incident bridges when required.Communicate clearly with engineering, product, and operations teams during critical incidents.Analyze logs, metrics, alerts, and system behavior to proactively identify potential issues.Improve monitoring and alerting to reduce noise and increase actionable signals.Identify opportunities to improve system reliability, resilience, and availability.Automate repetitive operational tasks and reduce manual effort.Develop and improve operational tooling for incident response, monitoring, and production support.Troubleshoot infrastructure across on-premises, cloud, and containerized environments.Diagnose issues involving servers, networking, Kubernetes, cloud infrastructure, and application dependencies.Collaborate with engineering teams to improve production stability and operational processes.
Required Technical Skills & Experience
Cloud & Infrastructure
8+ years of experience in SRE, Production Support, DevOps, Infrastructure Engineering, or a related field.Strong hands-on experience with AWS services such as:EC2S3LambdaLoad BalancersECSExperience with GCP and/or Azure.Experience supporting hybrid and/or multi-cloud environments.Strong Linux administration and troubleshooting skills.Experience with Windows and VMware environments is a plus.
Kubernetes & Containers
Hands-on experience troubleshooting Kubernetes environments.Strong understanding of pods, ingress, configurations, networking, and cluster-level issues.Understanding of containerized application architectures.Experience with Docker and related container technologies.
DevOps & CI/CD
Strong understanding of DevOps practices and CI/CD pipelines.Hands-on experience with Harness.Experience with GitHub and/or GitLab.Experience automating operational processes and workflows.
Application & Database Troubleshooting
Working knowledge of Java, Node.js, and React-based applications.Strong understanding of application dependencies and service-to-service communication.Experience troubleshooting database connectivity and application dependencies.Working knowledge of Oracle, MariaDB, and/or MSSQL environments.
Networking
Strong understanding of TCP/IP, DNS, HTTP, and HTTPS.Experience troubleshooting connectivity, latency, load balancing, and service-to-database communication.Understanding of networking concepts within cloud and Kubernetes environments.Experience with Akamai CDN and traffic-management technologies is a plus.
Incident Management & Observability
Strong experience with production incident management and triage.Experience troubleshooting high-impact production issues under pressure.Experience with monitoring, logging, alerting, and observability tools.Prior experience acting as an Incident Commander or leading major production incidents is highly preferred.Strong understanding of Root Cause Analysis (RCA) and problem management.
Preferred Qualifications
Experience supporting large-scale enterprise production environments.Experience working with highly available, high-volume applications.Experience in globally distributed engineering or operations teams.Prior Incident Commander experience.Strong automation and reliability engineering mindset.Ability to identify recurring issues and implement long-term solutions.Familiarity with AI-assisted development and automation tools is a plus.