Summary
✨ AI‑Generated
Work remotely as a senior SRE responsible for a modern cloud-native platform. You will build observability and alerting systems, define meaningful reliability objectives, improve incident response, optimize progressive delivery, automate operational work, and use performance engineering to strengthen scalability and release safety.
Highlights
Senior SRE role with hands-on ownership of observability, reliability, incident response, and safe deployments. Offers remote flexibility plus access to offices, with substantial work in AWS, Kubernetes, monitoring, performance engineering, and infrastructure automation.
Description
We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis.
Unlock the potential of remote work in Kazakhstan, giving you the flexibility to work from home or access our offices in Astana, Almaty or Karaganda.
Responsibilities
Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooksDefine meaningful SLIs/SLOs and reduce alert noise to improve signal qualityDesign and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanismsApply a performance engineering mindset through load testing, capacity analysis, and latency profilingAutomate operational toil through scripting and infrastructure-as-codeAccelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effortLead incident response practices including on-call readiness and blameless post-mortemsCollaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements
Requirements
3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systemsExpertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and eventsKnowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviewsExperience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layersProficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelinesBackground in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotationsA proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teamsStrong communication skills to turn noisy signals into crisp findings, runbooks, and recommendationsEnglish Level: B2+ (Upper-Intermediate) or higher
Nice to have
Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactivelySkills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecksFamiliarity with leveraging AIOps capabilities to advance SRE maturity and drive innovationExperience with Datadog or similar enterprise observability platformsBackground in evangelizing best practices and setting standards across engineering teamsExposure to programmatic advertising or adtech platforms
We offer
We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersDiscounts in local language schools, including online courses for the Kazakh languageWe cover it all:Participation in the Employee Stock Purchase PlanMonetary bonuses for engaging in the referral programMedical & family care packageSix trust days per year (sick leave without a medical certificate)Coverage of psychology sessions of your choiceBenefits package (sports activities, a variety of stores and services)Housing support program: preferential mortgage access via Otbasy Bank partnership
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups.
With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.