Summary
✨ AI‑Generated
A senior ML infrastructure role for an engineer who enjoys operating production systems. You will own the infrastructure, deployment, and operations required to run real-time inference services at scale, including monitoring, incident response, debugging, rollbacks, and postmortems. Proven experience with 24x7 on-call rotations for high-load production services is required. Remote work is available from within the country.
Highlights
Work with applied scientists and own production infrastructure end to end for large-scale machine learning services. The role provides genuine operational ownership, remote flexibility within the country, and direct impact on reliability and production performance.
Description
We are looking for a Senior ML Infrastructure Engineer who enjoys running systems in production.
You will work side by side with our applied scientists: they build the models, and you own everything needed to run them reliably for millions of customers.
Our team owns its infrastructure end to end, deployment and operations included, so your work is visible, and the ownership is real.
Experience with 24x7 on-call rotations for high-load online services in production is required for this role.
You will thrive here if you have been the person who gets paged and fixes things, rather than handing deployment and operations to a separate team.
Experience the freedom of remote work from anywhere in Georgia, whether from the comfort of your home, our modern offices in Tbilisi and Batumi or a coworking space in Kutaisi.
Responsibilities
Keep real-time inference services healthy through alerting, live debugging and fixing of production incidents, rollbacks and postmortems, as part of a shared 24x7 on-call rotationBuild and improve AWS infrastructure, including infrastructure as code, CI/CD pipelines, Kubernetes, Docker, autoscaling, monitoring, load testing and cost controlCare for the GPU fleet through NVIDIA driver and CUDA upgrades, node provisioning and debugging, and capacity managementAutomate user access, including SSH credentials, service accounts and developer environments for the science teamRun real-time production data pipelines, including streaming ingestion into an online feature store and serving features to low-latency endpointsOrchestrate batch jobs with Airflow or Databricks WorkflowsPartner with applied scientists to productionize their prototypes and keep them healthyWrite and review production Python, including validating AI-generated code
Requirements
24x7 on-call experience with high-load online services in production, including real incidents personally debugged and fixed liveStrong Linux systems skills, with hands-on experience in Kubernetes, Docker and infrastructure as codeSolid experience in Python, Git and CI/CDExperience with workflow orchestration tools such as Airflow, Databricks Workflows or equivalentProficiency in Apache Airflow, CI/CD and DevOps practicesFamiliarity with Gen AI in SDLCEnglish proficiency at B2 level or higher
Nice to have
Familiarity with Anthropic Claude Code
We offer
We connect like-minded peopleDelivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growthLeadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersWe cover it allParticipation in the Employee Stock Purchase PlanMonetary bonuses for engaging in the referral programComprehensive medical & family care packageFive trust days per year (sick leave without a medical certificate)Benefits package (sports activities, a variety of stores and services)
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups.
With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.