Site Reliability Engineer

Livemindz โ€” United States ยท Posted ~22 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Required qualifications Core ResponsibilitiesPlatform Reliability: Manage service availability, resilience, and operational efficiency across large-scale microservices architecture.Observability & Monitoring: Administer and optimize stacks including Prometheus, Grafana, AlertManager, and OpenSearch.Automation & CI/CD: Write automation scripts in Python and build pipelines via GitHub Actions.Incident Management: Lead tier-3 production support, root-cause analysis, and systematic toil reduction . 8+ years experience with site reliability engineering practices, including monitoring, incident response, and system performance optimization.Proficiency in Python, Docker and EKS.Familiarity with cloud platforms such as AWS, Azure, or Google Cloud.Experience with containerization and orchestration tools like Docker and Kubernetes.Knowledge of infrastructure as code tools such as Terraform or Ansible.Experience with CI/CD pipelines and automation tools.Strong understanding of Linux/Unix systems and networking concept s.