Senior AI DevOps/SRE Engineer

Epam Systems — Uzbekistan · Posted ~2 days ago

Senior Full-time Remote

Skills

AI DevOps SRE CI/CD MLOps AI deployment machine learning operations LLM deployment RAG systems monitoring reliability engineering DevOps LLMs RAG machine learning

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior AI DevOps/SRE engineer is sought to ensure reliable deployment, operation, and scaling of AI and machine learning systems, including LLM and retrieval-augmented generation workloads. Responsibilities include CI/CD automation, monitoring, reliability engineering, deployment strategy, and collaboration with data science and software engineering teams.

Highlights

Senior remote role focused on deploying, scaling, and operating advanced AI systems. The position combines established DevOps and SRE practices with MLOps, CI/CD, reliability, and AI infrastructure, while working closely with data scientists and software engineers on high-impact AI initiatives.

Description

We are currently seeking an experienced Senior AI DevOps/SRE to join our team. In this pivotal role, you will collaborate closely with data scientists and software developers to ensure seamless integration and optimize the operational efficiency of our AI deployments. Your expertise will be pivotal in deploying, maintaining, and scaling our cutting-edge AI solutions, encompassing LLMs and RAG systems. As a key team member, you will spearhead both traditional DevOps responsibilities and innovative approaches to MLOps. Your proactive involvement will be essential in driving the success of our AI initiatives and maximizing their impact across the organization. Experience the freedom of remote work from anywhere in Uzbekistan, whether it's the comfort of your home or our modern office in Tashkent. Responsibilities Implement and maintain CI/CD pipelines for AI and machine learning projects, ensuring robust deployment strategies and continuous integrationMonitor and ensure the reliability, availability, and performance of AI applications, particularly those involving LLMs and RAGCollaborate with AI research teams to operationalize machine learning models and systems efficientlyDevelop and enforce best practices for version control, configuration management, and testing of AI-driven software solutionsUtilize MLOps tools such as Kubeflow, MLflow, or TensorFlow Extended (TFX) to streamline the machine learning lifecycle from experimentation to productionImplement monitoring solutions that track both system metrics and model performance to facilitate proactive issue resolutionParticipate in on-call rotations to support the operational health of critical systems, employing SRE principles to meet service-level objectives (SLOs) and reduce downtime Requirements Bachelor’s degree in Computer Science, Engineering, or a related fieldProven experience as a DevOps Engineer or SRE, with a strong background in software development and automationExpertise in deployment and management of LLMs, including technologies like RAGProficient in CI/CD tools (Jenkins, GitLab CI, CircleCI) and infrastructure as code (Terraform, Ansible)Solid knowledge of container orchestration technologies (Kubernetes, Docker)Familiarity with MLOps tools and practices to support machine learning lifecycle management Nice to have Experience with cloud services (AWS, GCP, Azure), particularly in AI/ML deploymentsBackground in monitoring tools like Prometheus, Grafana, and ELK stackUnderstanding of Python, particularly in data science and machine learning contextsCertification in Kubernetes, AWS/GCP/Azure, or similar technologies We offer We connect like-minded people:Delivering innovative solutions to industry leaders, making a global impactEnjoyable working environment, whether it is the vibrant office or the comfort of your own homeOpportunity to work abroad for up to two months per yearRelocation opportunities within our offices in 55+ countriesCorporate and social eventsWe invest in your growth:Leadership development, career advising, soft skills and well-being programsCertifications, including GCP, Azure and AWSUnlimited access to EPAM's internal learning databaseFree English classes with certified teachersDiscounts in local language schools, including offline courses for the Uzbek languageWe cover it all:Monetary bonuses for engaging in the referral programMedical & family care packageFour trust days per year (sick leave without a medical certificate)Discounts for fitness clubs, dance schools and sports programsBenefits package (sports activities, a variety of stores and services) EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments. Not found