Site Reliability Engineer

Sotalentjobs — United States · Posted ~2 hours ago

Senior Remote

Skills

SRE distributed systems automation observability incident response

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an experienced Site Reliability Engineer to improve system reliability, automate operations, manage incidents, and enhance large-scale distributed platforms.

Highlights

Remote role focused on reliability engineering, automation, and scalable systems. Provides opportunities to improve performance and operational excellence.

Description

Site Reliability Engineer 📍 Location: United States (Remote) 🏢 Industry: IT System Custom Software Development 💼 Work Setting: Remote Are you passionate about building highly reliable, scalable, and resilient systems that power critical business operations? We are seeking an experienced Site Reliability Engineer (SRE) to drive platform reliability, operational excellence, and automation across large-scale distributed environments. In this role, you will bridge software engineering and infrastructure operations, leveraging automation, observability, and engineering best practices to enhance system performance, reduce operational overhead, and ensure exceptional service reliability. Key Responsibilities Define, measure, and continuously improve service reliability through service-level objectives (SLOs), service-level indicators (SLIs), and performance metrics.Lead the response, coordination, and resolution of production incidents while facilitating root cause analysis and continuous improvement initiatives.Design, implement, and maintain comprehensive monitoring, logging, alerting, and observability solutions to provide actionable insights into system health and performance.Develop and enhance operational processes, runbooks, escalation procedures, and support frameworks to improve service availability and incident response effectiveness.Automate repetitive operational tasks and workflows through scripting, tooling, and infrastructure automation.Manage and optimize containerized platforms and orchestration environments, including scalability, availability, networking, and performance considerations.Design and support modern CI/CD practices that enable secure, reliable, and efficient software delivery.Conduct capacity planning, performance analysis, stress testing, and resilience validation to support current and future business growth.Collaborate closely with development teams to embed reliability, scalability, and operational readiness throughout the software development lifecycle.Strengthen platform resilience through fault tolerance, high-availability architectures, recovery strategies, and proactive risk mitigation practices.Partner with security teams to support vulnerability management, system hardening, patching, and secure infrastructure operations.Contribute to the roadmap for platform engineering, observability, automation, and developer productivity improvements.Mentor team members and promote a culture of continuous learning, operational excellence, collaboration, and accountability.Required Qualifications Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field.10+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Production Operations, or related disciplines.Strong programming and automation skills using languages such as Python, Go, Java, or equivalent technologies.Extensive experience administering and troubleshooting Linux-based environments.Hands-on experience managing containerized applications and orchestration platforms in production environments.Strong understanding of monitoring, observability, logging, and performance management practices.Experience designing, implementing, and supporting CI/CD pipelines and automated deployment processes.Solid knowledge of distributed systems, system architecture, scalability, and reliability engineering principles.Proven experience leading production incident management and post-incident review activities.Excellent analytical, problem-solving, communication, and documentation skills.Preferred Qualifications Experience implementing and managing reliability frameworks, service performance metrics, and operational excellence programs.Knowledge of resilience testing, fault injection, or reliability validation methodologies.Hands-on experience with major cloud platforms and modern infrastructure technologies.Background in performance engineering, capacity planning, scalability analysis, or load testing.Familiarity with service networking, platform modernization, and cloud-native technologies.Experience influencing engineering best practices and driving cross-functional reliability initiatives.