Senior Site Reliability Engineer

Infojini Inc — Chile · Posted ~4 days ago

Senior Full-time Remote

Skills

Site Reliability Engineering Production incident response Distributed systems troubleshooting Root cause analysis Automation Observability Production operations Incident management Technical communication Distributed systems Production systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior Site Reliability Engineering role supporting large-scale enterprise production systems in a highly available distributed environment. You’ll respond to critical incidents, troubleshoot complex systems, lead major incident bridges when needed, improve observability and automation, and drive continuous reliability improvements. The role follows an evening-to-overnight Eastern Time schedule and is remote within Latin America.

Highlights

Fully remote senior SRE work supporting large-scale, highly available enterprise systems, with hands-on ownership of reliability, automation, observability and major incident response.

Description

Senior Site Reliability Engineer (Sr. SRE) Location: LATAM – Remote Work Hours: 5:00 PM – 1:00 AM EST, Experience: 8+ years About the Role We are seeking a Senior Site Reliability Engineer (SRE) to support large-scale enterprise production applications in a highly available and distributed environment. This is a hands-on role focused on production incident response, troubleshooting, reliability engineering, automation, observability, and continuous improvement. The ideal candidate will have strong technical troubleshooting skills, excellent communication abilities, and experience working in high-pressure production environments. Key Responsibilities Respond to production alerts and incidents, assess severity, and initiate appropriate mitigation actions.Troubleshoot complex distributed systems and identify root causes of production issues.Participate in and lead major incident bridges when required.Communicate clearly with engineering, product, and operations teams during critical incidents.Analyze logs, metrics, alerts, and system behavior to proactively identify potential issues.Improve monitoring and alerting to reduce noise and increase actionable signals.Identify opportunities to improve system reliability, resilience, and availability.Automate repetitive operational tasks and reduce manual effort.Develop and improve operational tooling for incident response, monitoring, and production support.Troubleshoot infrastructure across on-premises, cloud, and containerized environments.Diagnose issues involving servers, networking, Kubernetes, cloud infrastructure, and application dependencies.Collaborate with engineering teams to improve production stability and operational processes. Required Technical Skills & Experience Cloud & Infrastructure 8+ years of experience in SRE, Production Support, DevOps, Infrastructure Engineering, or a related field.Strong hands-on experience with AWS services such as:EC2S3LambdaLoad BalancersECSExperience with GCP and/or Azure.Experience supporting hybrid and/or multi-cloud environments.Strong Linux administration and troubleshooting skills.Experience with Windows and VMware environments is a plus. Kubernetes & Containers Hands-on experience troubleshooting Kubernetes environments.Strong understanding of pods, ingress, configurations, networking, and cluster-level issues.Understanding of containerized application architectures.Experience with Docker and related container technologies. DevOps & CI/CD Strong understanding of DevOps practices and CI/CD pipelines.Hands-on experience with Harness.Experience with GitHub and/or GitLab.Experience automating operational processes and workflows. Application & Database Troubleshooting Working knowledge of Java, Node.js, and React-based applications.Strong understanding of application dependencies and service-to-service communication.Experience troubleshooting database connectivity and application dependencies.Working knowledge of Oracle, MariaDB, and/or MSSQL environments. Networking Strong understanding of TCP/IP, DNS, HTTP, and HTTPS.Experience troubleshooting connectivity, latency, load balancing, and service-to-database communication.Understanding of networking concepts within cloud and Kubernetes environments.Experience with Akamai CDN and traffic-management technologies is a plus. Incident Management & Observability Strong experience with production incident management and triage.Experience troubleshooting high-impact production issues under pressure.Experience with monitoring, logging, alerting, and observability tools.Prior experience acting as an Incident Commander or leading major production incidents is highly preferred.Strong understanding of Root Cause Analysis (RCA) and problem management. Preferred Qualifications Experience supporting large-scale enterprise production environments.Experience working with highly available, high-volume applications.Experience in globally distributed engineering or operations teams.Prior Incident Commander experience.Strong automation and reliability engineering mindset.Ability to identify recurring issues and implement long-term solutions.Familiarity with AI-assisted development and automation tools is a plus.