Senior Site Reliability Engineer

Epam Systems — Uzbekistan · Posted ~2 days ago

Senior Full-time Remote

Skills

Site Reliability Engineering Grafana AWS Kubernetes EKS SLIs/SLOs monitoring alerting incident response CI/CD progressive delivery infrastructure as code load testing Amazon EKS Infrastructure as Code

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior SRE will own observability and reliability improvements for a cloud-native platform. Responsibilities include building dashboards and runbooks, defining meaningful SLIs/SLOs, reducing alert noise, improving deployment safety, conducting performance analysis, and automating operational work. The role is fully remote within Uzbekistan.

Highlights

Fully remote senior SRE opportunity with strong ownership of observability, reliability engineering, deployment safety, performance, and operational automation.

Description

We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis. Experience the freedom of remote work from anywhere in Uzbekistan, whether it's the comfort of your home or our modern office in Tashkent. Responsibilities Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooksDefine meaningful SLIs/SLOs and reduce alert noise to improve signal qualityDesign and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanismsApply a performance engineering mindset through load testing, capacity analysis, and latency profilingAutomate operational toil through scripting and infrastructure-as-codeAccelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effortLead incident response practices including on-call readiness and blameless post-mortemsCollaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements Requirements 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systemsExpertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and eventsKnowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviewsExperience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layersProficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelinesBackground in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotationsA proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teamsStrong communication skills to turn noisy signals into crisp findings, runbooks, and recommendationsEnglish Level: B2+ (Upper-Intermediate) or higher Nice to have Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactivelySkills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecksFamiliarity with leveraging AIOps capabilities to advance SRE maturity and drive innovationExperience with Datadog or similar enterprise observability platformsBackground in evangelizing best practices and setting standards across engineering teamsExposure to programmatic advertising or adtech platforms We offer We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersDiscounts in local language schools, including offline courses for the Uzbek languageWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageFour trust days per year (sick leave without a medical certificate)Discounts for fitness clubs, dance schools and sports programsBenefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.