Senior Site Reliability Engineer

Arkhyatech — United States · Posted ~4 hours ago

Senior Hybrid

Skills

DevOps AWS Infrastructure as Code Terraform Kubernetes Docker Podman Linux Networking CI/CD Splunk Grafana SSL/TLS

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering opportunity for an experienced DevOps and platform specialist. You will design, operate, and optimize cloud and on-premises infrastructure, build resilient container platforms, strengthen CI/CD pipelines, and improve observability, automated validation, and rollback capabilities. Strong hands-on experience with AWS, infrastructure as code, Kubernetes, Linux, networking, monitoring, and security certificates is required.

Highlights

Senior platform role focused on cloud infrastructure, reliability engineering, automation, observability, and large-scale containerized systems.

Description

Senior Site Reliability Engineer (SRE) – DevOps & Platform Location: San Francisco, CA 94105 (Hybrid) Experience: 10–12 Years Key Responsibilities Demonstrate strong hands-on expertise in DevOps, cloud infrastructure, and platform engineering.Design, deploy, manage, and optimize AWS cloud infrastructure, with strong experience in Infrastructure as Code (IaC), preferably Terraform.Design, deploy, and operate containerized applications using Kubernetes and container technologies such as Docker and Podman across on-premises and cloud environments.Provide strong administration and troubleshooting expertise across Linux systems, networking, and CI/CD pipelines.Build reliability into deployment pipelines through automated health checks, observability, validation, and rollback mechanisms.Implement and maintain monitoring and observability solutions using tools such as Splunk and Grafana.Manage SSL/TLS certificates, including certificate lifecycle management, SSL handshakes, multi-domain certificates, and certificate management within Kubernetes environments.Partner closely with release-focused teams and application engineering teams to support reliable, secure, and efficient end-to-end application delivery.Apply SRE principles and best practices to improve system reliability, availability, scalability, performance, and operational efficiency.Participate in incident response, root-cause analysis, capacity planning, and continuous improvement of platform and production environments.Nice to Have Experience with Python scripting for automation, tooling, and platform engineering.Knowledge or hands-on experience with Honeycomb and modern observability practices.