Senior Site Reliability Engineer

Epam Systems — Uzbekistan · Posted ~14 hours ago

Senior Full-time Remote

Skills

SRE AWS Kubernetes observability SLI/SLO incident response infrastructure as code Grafana CI/CD IaC

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior reliability engineering role responsible for improving system stability through monitoring, automation, cloud infrastructure, deployment safety, and incident management practices.

Highlights

Remote senior SRE role focused on reliability engineering, automation, observability, and scalable cloud platforms.

Description

We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis. Experience the freedom of remote work from anywhere in Uzbekistan, whether it's the comfort of your home or our modern office in Tashkent. Responsibilities Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooksDefine meaningful SLIs/SLOs and reduce alert noise to improve signal qualityDesign and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanismsApply a performance engineering mindset through load testing, capacity analysis, and latency profilingAutomate operational toil through scripting and infrastructure-as-codeAccelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effortLead incident response practices including on-call readiness and blameless post-mortemsCollaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements Requirements 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systemsExpertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and eventsKnowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviewsExperience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layersProficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelinesBackground in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotationsA proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teamsStrong communication skills to turn noisy signals into crisp findings, runbooks, and recommendationsEnglish Level: B2+ (Upper-Intermediate) or higher Nice to have Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactivelySkills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecksFamiliarity with leveraging AIOps capabilities to advance SRE maturity and drive innovationExperience with Datadog or similar enterprise observability platformsBackground in evangelizing best practices and setting standards across engineering teamsExposure to programmatic advertising or adtech platforms We offer We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersDiscounts in local language schools, including offline courses for the Uzbek languageWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageFour trust days per year (sick leave without a medical certificate)Discounts for fitness clubs, dance schools and sports programsBenefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.