Site Reliability Engineer

Vaspire Technologies Inc — United States · Posted ~4 hours ago

Mid Full-time

Skills

Infrastructure as Code CI/CD reliability engineering SLIs/SLOs IAM security compliance monitoring observability incident management cloud cost optimization Terraform GitHub Actions ArgoCD metrics logs distributed tracing cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is looking for a Site Reliability Engineer to automate infrastructure and deployments, establish reliability metrics, strengthen cloud security and compliance, and improve observability. The role includes incident response, on-call operations, post-mortems, cloud cost optimization, documentation, and mentoring.

Highlights

Own reliability, security, observability, incident response, and cloud cost optimization while improving infrastructure automation and engineering practices.

Description

As a Site Reliability Engineer, you will: Infrastructure as Code & CI/CD: Automate provisioning and deployments with Terraform and integrate best-practice pipelines (GitHub Actions, ArgoCD, etc.).Reliability Engineering: Define SLIs/SLOs, manage error budgets, and build dashboards & alerts to proactively measure and improve system health.Security & Compliance: Enforce least-privilege IAM policies, automate vulnerability scans, and maintain audit logging for compliance.Monitoring & Observability: Instrument services with metrics, logs, and distributed tracing to enable rapid troubleshooting, aid teams in alerting, custom metrics, and dashboarding.Incident Management: Own on-call rotations, lead real-time incident response, conduct post-mortems, and drive continuous improvements.Cost Optimization: Implement tagging strategies, right-size resources, and leverage concrete data to decide on optimal methods to control cloud spend at scale.Documentation & Mentorship: Author runbooks, standards, and best-practice guides—and coach dev teams on implementing modern DevOps, reliability, and security patterns. Qualifications Have 5+ years of experience running production critical systemsDeep proficiency with the AWS Cloud and Cloud-Native best practicesExperience with Kubernetes (EKS, GKE) and Container Orchestration at scaleSkilled in Terraform to declaratively provision and maintain infrastructure servicesWorking knowledge of managing and debugging databases like Redis and PostgresStrong familiarity with VPC, VPN, Load Balancing, and cloud networking componentsProficiency with Git workflows, branching strategies, and CI/CD system integrationsSolid understanding of web and network protocols and standards (HTTP, REST, TLS, DNS, etc...) Required Skills Professional proficiency in English (both written and spoken) is required for this role. Preferred Skills Bachelor's degree, or equivalent in Computer Science, Engineering, or a related field.Experience with ArgoCD, Github Actions, Jenkins, or other CI/CD pipeline solutionsWorking knowledge of Python, Golang, and Helm templating languagesNode.js experience a plus, including running scalable, resilient Node microservicesGrasp of foundational security best practices for cloud infrastructureAwareness of Terragrunt, managing Terraform state, and optimal project structureSeasoned in production readiness fundamentals amidst a fast moving team