Principal Site Reliability Engineer

Epam Systems — Kyrgyzstan · Posted ~10 hours ago

Lead Full-time

Skills

Site Reliability Engineering AWS Kubernetes Terraform GitOps Cloud infrastructure Incident response Kotlin React Native

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior individual contributor role focused on building and operating cloud-native infrastructure platforms. The engineer will design automation, deployment workflows, reliability practices, and self-service capabilities for development teams.

Highlights

Lead the design of a modern cloud-native platform from the ground up, creating developer enablement tools and reliability frameworks with significant technical ownership.

Description

We are seeking an experienced Principal Site Reliability Engineer (SRE) to architect, build, and operate the foundational platform infrastructure for a greenfield, cloud-native platform on AWS. This role is 100% focused on proactive platform engineering, developer enablement, and reliability architecture, not daily ticket handling or manual operations. Operating as an individual contributor, you will design and implement an Internal Developer Platform (IDP) to empower Kotlin backend and React Native mobile development teams, while establishing the organization's incident response frameworks, on-call models, and observability standards from scratch. Responsibilities Build self-service developer tooling, golden paths, and automated environment provisioning pipelines so development teams can deploy microservices safely and independentlyProvision, harden, and manage production-grade Amazon EKS clusters using modular Terraform, Karpenter autoscaling, and GitOps delivery patternsDesign and establish the organization's incident response model, on-call escalation policies, and blameless post-mortem processes from the ground upLead the enterprise implementation of Datadog, including APM, distributed tracing, and custom metrics, while defining meaningful Service Level Objectives (SLOs) and actionable alerting rulesArchitect GitHub Actions CI/CD workflows supporting zero-downtime progressive delivery strategies such as Canary and Blue/Green releases, along with automated health verificationPartner with the Solution Architect to implement robust cloud network segregation, including VPCs, transit gateways, IRSA, and ingress security boundaries Requirements 7+ years of experience as a Site Reliability Engineer, Platform Engineer, or similar role focused on cloud-native infrastructureExpertise in AWS, Amazon EKS, and Kubernetes in production environmentsProficiency in Terraform, Karpenter, and GitOps delivery patternsBackground in designing incident response frameworks, on-call models, and blameless post-mortem processesSkills in Datadog for APM, distributed tracing, and custom metrics, along with defining SLOs and alerting rulesCompetency in GitHub Actions for CI/CD workflows and progressive delivery strategies such as Canary and Blue/Green releasesUnderstanding of cloud network architecture, including VPCs, transit gateways, IRSA, and ingress security boundariesFamiliarity with Kotlin backend and React Native mobile development ecosystems to effectively support developer enablementEnglish proficiency at B2 level or higher We offer We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageSix trust days per year (sick leave without a medical certificate)Coverage of psychology sessions of your choiceDiscounts for fitness clubs and sports programsBenefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments. Experience the freedom of remote work from anywhere in Kyrgyzstan, whether it's the comfort of your home or our modern office in Bishkek.