Site Reliability Engineer

Insight Global — United States · Posted ~4 hours ago

Senior Contract Hybrid $175K-$200K

Skills

SRE Infrastructure Operations Monitoring Observability SLI/SLO Incident Response Automation Infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A reliability engineering role responsible for improving production systems through automation, monitoring, performance optimization, incident management, and infrastructure best practices.

Highlights

High-impact reliability engineering role focused on system availability, automation, resilience, and production operations.

Description

Title: Site Reliability Engineer Location: Arlington, VA Hybrid: 3 days on-site / 2 days remote Contract: 6 month to perm Clearance: Secret Salary: $175K-$200K Our client is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is based in Arlington, VA, as a hybrid/remote position. Responsibilities: Design and maintain highly available production systems.Define and manage SLIs, SLOs, and error budgets.Automate operational tasks and eliminate manual processes.Develop monitoring, alerting, and observability solutions.Improve system performance, capacity, and resilience.Lead incident response and root cause analysis.Implement disaster recovery and continuity strategies.Partner with development teams to improve application reliability. Required Skills and Experience: Bachelor's with 12+ years of infrastructure/cloud engineering experience (or commensurate experience)5–10+ years of engineering experience, with a strong background in Linux and Windows systemsExpertise in Kubernetes and container platformsExperience working with cloud infrastructure environmentsProficiency in scripting languages such as Python and GoHands-on knowledge of Terraform and automation toolsFamiliarity with monitoring platforms and incident management practicesExperience designing and managing CI/CD pipelines Preferred Skills and Experience: Kubernetes certificationsAWS/Azure certificationsDevOps certificationsITIL preferred