Summary
✨ AI‑Generated
We are seeking a Principal Site Reliability Engineer to architect and operate foundational platform infrastructure for a cloud-native platform on AWS. You will design an Internal Developer Platform, establish incident response frameworks, on-call models, and observability standards, and build self-service tooling for development teams.
Highlights
Opportunity to build cloud-native platform infrastructure from scratch, shape incident response, and enable development teams with self-service tooling.
Description
We are seeking an experienced Principal Site Reliability Engineer (SRE) to architect, build, and operate the foundational platform infrastructure for a greenfield, cloud-native platform on AWS.
This role is 100% focused on proactive platform engineering, developer enablement, and reliability architecture, not daily ticket handling or manual operations.
Operating as an individual contributor, you will design and implement an Internal Developer Platform (IDP) to empower Kotlin backend and React Native mobile development teams, while establishing the organization's incident response frameworks, on-call models, and observability standards from scratch.
Unlock the potential of remote work in Kazakhstan, giving you the flexibility to work from home or access our offices in Astana, Almaty or Karaganda.
Responsibilities
Build self-service developer tooling, golden paths, and automated environment provisioning pipelines so development teams can deploy microservices safely and independentlyProvision, harden, and manage production-grade Amazon EKS clusters using modular Terraform, Karpenter autoscaling, and GitOps delivery patternsDesign and establish the organization's incident response model, on-call escalation policies, and blameless post-mortem processes from the ground upLead the enterprise implementation of Datadog, including APM, distributed tracing, and custom metrics, while defining meaningful Service Level Objectives (SLOs) and actionable alerting rulesArchitect GitHub Actions CI/CD workflows supporting zero-downtime progressive delivery strategies such as Canary and Blue/Green releases, along with automated health verificationPartner with the Solution Architect to implement robust cloud network segregation, including VPCs, transit gateways, IRSA, and ingress security boundaries
Requirements
7+ years of experience as a Site Reliability Engineer, Platform Engineer, or similar role focused on cloud-native infrastructureExpertise in AWS, Amazon EKS, and Kubernetes in production environmentsProficiency in Terraform, Karpenter, and GitOps delivery patternsBackground in designing incident response frameworks, on-call models, and blameless post-mortem processesSkills in Datadog for APM, distributed tracing, and custom metrics, along with defining SLOs and alerting rulesCompetency in GitHub Actions for CI/CD workflows and progressive delivery strategies such as Canary and Blue/Green releasesUnderstanding of cloud network architecture, including VPCs, transit gateways, IRSA, and ingress security boundariesFamiliarity with Kotlin backend and React Native mobile development ecosystems to effectively support developer enablementEnglish proficiency at B2 level or higher
We offer
We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersDiscounts in local language schools, including online courses for the Kazakh languageWe cover it all:Participation in the Employee Stock Purchase PlanMonetary bonuses for engaging in the referral programMedical & family care packageSix trust days per year (sick leave without a medical certificate)Coverage of psychology sessions of your choiceBenefits package (sports activities, a variety of stores and services)Housing support program: preferential mortgage access via Otbasy Bank partnership
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups.
With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.