Summary
✨ AI‑Generated
A Principal SRE is sought to architect and build the foundational infrastructure for a greenfield cloud-native platform. This individual-contributor role emphasizes proactive platform engineering rather than ticket handling, with responsibility for self-service developer tooling, automated environment provisioning, reliability architecture, incident response, on-call practices, and observability. The role is fully remote within the candidate's country and provides substantial ownership from the ground up.
Highlights
Principal-level individual contributor role focused entirely on proactive platform engineering, reliability architecture, developer enablement, and automation. Offers full remote work within Uzbekistan and the opportunity to build an internal developer platform, incident-response practices, on-call models, and observability standards from the ground up.
Description
We are seeking an experienced Principal Site Reliability Engineer (SRE) to architect, build, and operate the foundational platform infrastructure for a greenfield, cloud-native platform on AWS.
This role is 100% focused on proactive platform engineering, developer enablement, and reliability architecture, not daily ticket handling or manual operations.
Operating as an individual contributor, you will design and implement an Internal Developer Platform (IDP) to empower Kotlin backend and React Native mobile development teams, while establishing the organization's incident response frameworks, on-call models, and observability standards from scratch.
Experience the freedom of remote work from anywhere in Uzbekistan, whether it's the comfort of your home or our modern office in Tashkent.
Responsibilities
Build self-service developer tooling, golden paths, and automated environment provisioning pipelines so development teams can deploy microservices safely and independentlyProvision, harden, and manage production-grade Amazon EKS clusters using modular Terraform, Karpenter autoscaling, and GitOps delivery patternsDesign and establish the organization's incident response model, on-call escalation policies, and blameless post-mortem processes from the ground upLead the enterprise implementation of Datadog, including APM, distributed tracing, and custom metrics, while defining meaningful Service Level Objectives (SLOs) and actionable alerting rulesArchitect GitHub Actions CI/CD workflows supporting zero-downtime progressive delivery strategies such as Canary and Blue/Green releases, along with automated health verificationPartner with the Solution Architect to implement robust cloud network segregation, including VPCs, transit gateways, IRSA, and ingress security boundaries
Requirements
7+ years of experience as a Site Reliability Engineer, Platform Engineer, or similar role focused on cloud-native infrastructureExpertise in AWS, Amazon EKS, and Kubernetes in production environmentsProficiency in Terraform, Karpenter, and GitOps delivery patternsBackground in designing incident response frameworks, on-call models, and blameless post-mortem processesSkills in Datadog for APM, distributed tracing, and custom metrics, along with defining SLOs and alerting rulesCompetency in GitHub Actions for CI/CD workflows and progressive delivery strategies such as Canary and Blue/Green releasesUnderstanding of cloud network architecture, including VPCs, transit gateways, IRSA, and ingress security boundariesFamiliarity with Kotlin backend and React Native mobile development ecosystems to effectively support developer enablementEnglish proficiency at B2 level or higher
We offer
We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersDiscounts in local language schools, including offline courses for the Uzbek languageWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageFour trust days per year (sick leave without a medical certificate)Discounts for fitness clubs, dance schools and sports programsBenefits package (sports activities, a variety of stores and services)
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups.
With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.