Site Reliability Engineer

Programmers Io — United States · Posted ~6 hours ago

Full-time Hybrid

Skills

Environment monitoring Task automation System troubleshooting SLO management SLA management Software and server configuration Database configuration

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an engineering team as a Site Reliability Engineer, improving the monitoring, automation, availability, and reliability of critical systems. You will optimize software, servers, databases, and service configurations, monitor SLOs and SLAs, troubleshoot live systems, and collaborate across technical and business teams.

Highlights

Hybrid work, ownership of reliability improvements, cross-functional collaboration, exposure to critical business systems, and opportunities to act as a technical subject matter expert.

Description

Job Role: SRE Engineer Location: Austin TX/South Lake, TX (Onsite- Hybrid) Employment Type: Full-Time Job Summary • Develop and maintain tooling used for environment monitoring and task automation • Identify application reliability and availability improvements and build solutions to drive an improved experience • Analyze and establish efficient configurations for software and servers, DB connections, indexes, drivers, etc. • Coordinate with development teams, technical and non-technical Partners and clients to maintain wide knowledge on dependencies of the critical business transaction including platform, services and tools • Monitor internal and vendor service level objectives (SLOs) and agreements (SLAs); identifies and resolves SLO / SLA gaps • Serve as technical subject matter expert (SME) for cross-functional engineering Teams; • Assist with and troubleshoot systems-related issues and maintenance • Collaborate on maintaining services once they are live; measures and monitors availability, latency, and overall system health • Develop run book and build automation • Develop and maintain E2E monitoring dashboards to support critical business transaction • Develop and maintain synthetic monitoring for critical business transaction using tools such as ThousandEyes • Practice sustainable incident response and blameless postmortems • Document and promote SRE standards and procedures • Develop and assist in deployment and rollback automation • Review Release and deployments requirements • Build and setup automation tests. • Incident communication to impacted stakeholders • Coach and mentor junior engineers and fellow practitioners Role Summary: Experienced SRE Engineer with 5+ years in designing, managing and supporting distributed systems across multi-cloud environments. Key Skills & Expertise: • CI/CD: GitHub, Harness • Cloud Platforms: GCP, PCF, AWS • Monitoring & Observability: Splunk, Grafana, AppDynamics, Thousand Eyes • Containers & Orchestration: Docker, Kubernetes, Cloud Foundry • Messaging & Streaming: Kafka, MQ • Protocols & Web Services: HTTP, DNS, TCP/UDP, REST, SOAP, JSON Core Competencies: • Strong troubleshooting and debugging in microservices architecture • Incident management, issue resolution and RCA creation • Multi-cloud platform management (SRE practices) • Enterprise cloud infrastructure handling • Agile development practices with tools like Git, Jira, Confluence