Site Reliability Engineer

Galenthq — United States · Posted ~6 hours ago

Senior Full-time Hybrid

Skills

SRE Linux Cloud platforms Docker Kubernetes Python Bash Prometheus Grafana Incident management Infrastructure as Code AWS Azure GCP

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology-focused organization is seeking an experienced reliability engineer to design scalable infrastructure, improve observability, automate operations, and enhance production stability using modern cloud-native practices.

Highlights

Opportunity to build highly available systems, improve reliability practices, automate operations, and work with modern cloud and container technologies.

Description

Job Title: SRE Engineer Location: Plano, TX (Hybrid, 3 Days Onsite) Job Summary Seeking an experienced Site Reliability Engineer (SRE) to design, build, and maintain highly available, scalable, and fault-tolerant systems. The role focuses on automation, observability, system reliability, incident management, and infrastructure modernization across cloud and containerized environments. Key Responsibilities Design and maintain highly available, scalable, and fault-tolerant systems.Implement and manage monitoring and observability tools such as Grafana and Prometheus.Participate in incident management, root cause analysis (RCA), and postmortem activities.Automate operational tasks using scripting and automation tools, including Python and Bash.Collaborate with development, infrastructure, and support teams to improve system reliability.Drive adoption of SRE practices, including SLIs, SLOs, and error budgets.Ensure performance optimization, capacity planning, and overall system stability.Build and maintain automation pipelines and Infrastructure as Code (IaC) solutions.Required Skills Strong experience in Site Reliability Engineering (SRE), Production Support, or DevOps environments.Hands-on experience with Linux/Unix systems.Experience with cloud platforms such as AWS, Azure, or GCP.Expertise with Docker and Kubernetes.Knowledge of monitoring and observability tools including Grafana, Prometheus, OpenSearch, and Instana.Strong troubleshooting and incident management skills.Experience with Infrastructure as Code tools such as Terraform and Ansible.Programming and scripting experience using Python, Bash, and PowerShell.Strong capabilities in Automation & Scripting.Experience with System Reliability Strategy Design.Preferred Skills Knowledge of Kafka and SQL/NoSQL databases.Understanding of ITIL processes and SRE concepts, including Toil and SLOs.Exposure to CI/CD tools such as Jenkins.Experience with multi-cloud or hybrid cloud environments.Qualifications Proven experience supporting and improving reliability, scalability, and operational excellence in complex technology environments.Ability to work effectively across development, infrastructure, and support teams.Strong analytical, troubleshooting, and automation skills. We are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, citizenship status, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law. https://www.e-verify.gov/sites/default/files/everify/posters/IER_RighttoWorkPoster.pdf