Site Reliability Engineer

Ipolarityllc โ€” United States ยท Posted ~1 day ago

Senior Full-time Onsite

Skills

SRE GCP Kubernetes Go Python Java Rust OpenTelemetry GraphQL Terraform CI/CD GKE Helm GitHub Actions Prometheus Grafana

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Support mission-critical applications in hybrid cloud environments by improving reliability, automation, observability, and performance using modern cloud-native technologies.

Highlights

Work on highly available enterprise systems with modern cloud infrastructure, observability, automation, and distributed systems technologies.

Description

Title: Site Reliability Engineer (SRE) Location: Scottsdale, AZ (Onsite) Employment Type: W2 Experience: 7+ Years Interview Process 1 Internal Evaluation 2 Client Telephonic Interviews 1 Final In-Person Interview (Richardson, TX or Scottsdale, AZ) We are seeking an experienced Site Reliability Engineer (SRE) to support and enhance large-scale, high-availability enterprise applications running across hybrid cloud environments . Required Skil ls7+ years of Site Reliability Engineering, Production Support, or DevOps experien ceExperience managing large-scale, high-performance applications in hybrid (on-premises and cloud) environmen tsStrong programming/scripting skills in Go, Python, Java, or Ru stExperience developing automation scripts and Application Performance Monitoring (APM) dashboar dsHands-on experience with Google Cloud Platform (GCP) and containerization technologi esStrong Kubernetes experience, including GKE/RKE/A KEExperience with OpenTelemetry (OTEL) for observability and distributed traci ngKnowledge of GraphQL frameworks such as Apollo, Prisma, or Hasu raExperience with databases such as Oracle, PostgreSQL, SQL Server, MongoDB, Redis, ClickHouse, or time-series databas esStrong understanding of networking concepts including TCP/IP, HTTP, DNS, Load Balancing, and Service Me shPreferred Skil lsMonitoring tools: Splunk, Grafana, Prometheus, AppDynamics, Dynatra ceInfrastructure as Code using Terraform and He lmCI/CD pipelines with GitHub Actio nsAutomation using Python, Ansible, and Node. jsHashiCorp Vault administrati onGoogle Cloud services including Cloud SQL, GCS, Spanner, Firestore, and BigQue ryExperience with Redis and in-memory cachi ngLinux/Windows administrati onVertex AI and Generative AI exposu reStrong troubleshooting skills for distributed systems and API Gateway environmen ts