Principal Site Reliability Engineer

Epam Systems — Kazakhstan · Posted ~2 hours ago

Lead Hybrid

Skills

Site Reliability Engineering Cloud Computing AWS Kubernetes Docker Observability Monitoring Incident Response On-call Infrastructure as Code Microservices Developer Enablement Terraform Prometheus Grafana CI/CD Python

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

We are seeking a Principal Site Reliability Engineer to architect and operate foundational platform infrastructure for a cloud-native platform on AWS. You will design an Internal Developer Platform, establish incident response frameworks, on-call models, and observability standards, and build self-service tooling for development teams.

Highlights

Opportunity to build cloud-native platform infrastructure from scratch, shape incident response, and enable development teams with self-service tooling.

Description

We are seeking an experienced Principal Site Reliability Engineer (SRE) to architect, build, and operate the foundational platform infrastructure for a greenfield, cloud-native platform on AWS. This role is 100% focused on proactive platform engineering, developer enablement, and reliability architecture, not daily ticket handling or manual operations. Operating as an individual contributor, you will design and implement an Internal Developer Platform (IDP) to empower Kotlin backend and React Native mobile development teams, while establishing the organization's incident response frameworks, on-call models, and observability standards from scratch. Unlock the potential of remote work in Kazakhstan, giving you the flexibility to work from home or access our offices in Astana, Almaty or Karaganda. Responsibilities Build self-service developer tooling, golden paths, and automated environment provisioning pipelines so development teams can deploy microservices safely and independentlyProvision, harden, and manage production-grade Amazon EKS clusters using modular Terraform, Karpenter autoscaling, and GitOps delivery patternsDesign and establish the organization's incident response model, on-call escalation policies, and blameless post-mortem processes from the ground upLead the enterprise implementation of Datadog, including APM, distributed tracing, and custom metrics, while defining meaningful Service Level Objectives (SLOs) and actionable alerting rulesArchitect GitHub Actions CI/CD workflows supporting zero-downtime progressive delivery strategies such as Canary and Blue/Green releases, along with automated health verificationPartner with the Solution Architect to implement robust cloud network segregation, including VPCs, transit gateways, IRSA, and ingress security boundaries Requirements 7+ years of experience as a Site Reliability Engineer, Platform Engineer, or similar role focused on cloud-native infrastructureExpertise in AWS, Amazon EKS, and Kubernetes in production environmentsProficiency in Terraform, Karpenter, and GitOps delivery patternsBackground in designing incident response frameworks, on-call models, and blameless post-mortem processesSkills in Datadog for APM, distributed tracing, and custom metrics, along with defining SLOs and alerting rulesCompetency in GitHub Actions for CI/CD workflows and progressive delivery strategies such as Canary and Blue/Green releasesUnderstanding of cloud network architecture, including VPCs, transit gateways, IRSA, and ingress security boundariesFamiliarity with Kotlin backend and React Native mobile development ecosystems to effectively support developer enablementEnglish proficiency at B2 level or higher We offer We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersDiscounts in local language schools, including online courses for the Kazakh languageWe cover it all:Participation in the Employee Stock Purchase PlanMonetary bonuses for engaging in the referral programMedical & family care packageSix trust days per year (sick leave without a medical certificate)Coverage of psychology sessions of your choiceBenefits package (sports activities, a variety of stores and services)Housing support program: preferential mortgage access via Otbasy Bank partnership EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.