Site Reliability Engineer

Pipercompanies — United States · Posted ~4 hours ago

Mid Contract Remote

Skills

Kubernetes Linux administration AWS cloud infrastructure automation observability alerting troubleshooting platform reliability security scalability Linux Cloud Automation Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A remote Site Reliability Engineering contract focused on operating and improving Kubernetes platforms across cloud and on-premises infrastructure. You will strengthen reliability, scalability, security, and operational performance through automation, monitoring, troubleshooting, and infrastructure engineering.

Highlights

Long-term remote contract focused on Kubernetes platform reliability across cloud and on-premises environments. Provides exposure to infrastructure automation, observability, security, scalability, and compliance-focused engineering in regulated environments.

Description

Piper Companies is seeking a Site Reliability Engineer (SRE) to support the development, maintenance, and operation of a Kubernetes-based platform within highly regulated cloud and on-premises environments. This individual will work closely with senior engineers and technical leaders to improve platform reliability, scalability, security, and operational performance while supporting compliance-driven initiatives. This is a long-tern contract opportunity with a strong focus on Kubernetes infrastructure, Linux administration, automation, and cloud technologies. This individual may sit remote in the US. Responsibilities for the Site Reliability Engineer include: Build, maintain, and support Kubernetes clusters across on-premises and AWS environments. Monitor platform reliability, availability, and performance through observability, alerting, and troubleshooting activities. Develop and implement automation tools and processes to improve operational efficiency and reduce manual intervention. Collaborate with senior engineers to define and track service reliability metrics, including SLIs, SLOs, and error budgets. Support compliance, security, auditing, and continuous monitoring initiatives within regulated environments. Contribute to Infrastructure as Code (IaC) development and enhancements using tools such as Terraform, CloudFormation, or similar technologies. Improve CI/CD pipelines and deployment processes to support platform scalability and operational excellence. Partner with Security, Platform, and Application teams to resolve issues and deliver reliable infrastructure solutions. Qualifications for the Site Reliability Engineer include: 4-6 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure-focused roles. Strong hands-on experience managing and supporting Kubernetes clusters in production environments. Experience with AWS, Azure, or similar cloud platforms; GovCloud experience is a plus. Strong Linux administration skills, including system troubleshooting, performance analysis, and infrastructure support. Experience with Infrastructure as Code technologies such as Terraform, CloudFormation, or similar tools. Programming or scripting experience in Python, Go, Ruby, or other object-oriented languages.. Experience with observability and monitoring tools such as Prometheus, Grafana, and centralized logging solutions. Exposure to CI/CD tools, deployment automation, and GitOps practices such as ArgoCD is preferred. Experience supporting FedRAMP High, DoD IL5, regulated, or audited environments is strongly preferred. Compensation for the Site Reliability Engineer includes: Salary range: $140,000 - $165,000 Comprehensive Benefits: Medical, Dental, Vision, 401(k), and applicable sick leave Keywords: Kubernetes, Site Reliability Engineering (SRE), DevOps, Platform Engineering, AWS, Azure, GovCloud, Linux, Terraform, CloudFormation, Infrastructure as Code (IaC), CI/CD, ArgoCD, Prometheus, Grafana, Monitoring, Alerting, Observability, Containerization, Docker, Networking, Automation, Python, Go, Ruby, Object-Oriented Programming, Cloud Infrastructure, Kubernetes Clusters, Production Support, Troubleshooting, Incident Response, Performance Optimization, Scalability, Reliability, High Availability, Security, Compliance, FedRAMP High, DoD IL5, Continuous Monitoring, Audit Support, GitOps, Deployment Automation, Platform Operations, Container Security, Cross-Functional Collaboration, SLI, SLO, Error Budgets, On-Call Support, Systems Administration, Infrastructure Management, Cloud-Native Technologies, AWS Infrastructure, Infrastructure Automation, Root Cause Analysis, Configuration Management, Platform Reliability, Regulated Environments. #REMOTE This job is open for applications on 9/25/2026 and will remain open for at least 30 days from the posting date