MLOps/AIOps Engineer
Extia — Romania · Posted ~1 day ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
📣 Would you like to join a company that puts people at the heart of its concerns? We are waiting for you! Since 2007, Extia, an IT consulting company, has been offering a unique approach in its field by combining well-being at work and performance.
Our philosophy at Extia is “First who, then what", so let's go for it!
🚀 First who?
You possess strong analytical and problem-solving skills, with the judgment to perfectly balance speed, reliability, and cost.You are a great communicator and collaborator, always ready to pave the way for data science, AI, and data engineering teams.
This role is hybrid, 2 days/week from the office in Bucharest.
🚀 Then what?
Senior MLOps / LLMOps / AIOps Engineer
You will be responsible for the entire process of bringing machine learning and AI models from the experimental phase to production, ensuring they remain reliable once running.
You will build the foundation upon which our data science and AI teams deliver results: automated training and deployment pipelines, model serving, and monitoring to detect model drift, latency, and runaway costs early.
Your tasks will include:
Designing and operating end-to-end ML pipelines—from data preparation to training, evaluation, and model registration—on Amazon SageMaker, Amazon MWAA (Managed Airflow), and EKS.Running production models on Amazon Bedrock and SageMaker, while building the capability to serve self-hosted models on internal infrastructure (containerized inference and appropriately sized GPU nodes).Owning the model lifecycle using a model registry, feature store, and MLflow for experiment tracking and versioning, including canary releases, automatic rollbacks, and clean promotion between environments.Building data and model monitoring to signal drift, performance degradation, and data quality issues, connecting these to automated retraining or alerts.Bringing AIOps into internal operations: using anomaly detection, event correlation, and root cause analysis based on metrics, logs, and traces to reduce alert noise and protect Service Level Objectives (SLOs).Operationalizing Generative AI (GenAI) solutions: deploying, tracking, and evaluating LLM workflows and agents on Amazon Bedrock, complete with guardrails, quality checks, and token/cost control.Managing the platform via Infrastructure as Code (IaC) with Terraform, maintaining secure, reproducible, and cost-effective environments.Raising the bar for production practices through code reviews and mentoring.Supporting AI governance through model lineage, reproducibility, and documentation aligned with the EU AI Act and GDPR.
Required technical skills:
Min.
5 years of experience in MLOps, DevOps, Platform Engineering, or SRE roles, including hands-on experience running ML systems in production.Solid programming skills in Python and shell scripting, alongside strong software engineering fundamentals (Git, testing, CI/CD).Solid hands-on experience with AWS (EKS / Kubernetes, Lambda, SageMaker, ECR, S3) in building and operating production systems.Strong competencies in Kubernetes, Docker, and IaC with Terraform.Experience building CI/CD pipelines in GitLab CI/CD and working with model lifecycle tools like MLflow and Apache Airflow / Amazon MWAA.Working knowledge of observability and monitoring (CloudWatch, Prometheus or VictoriaMetrics/Grafana, OTLP) and foundational AIOps concepts.Familiarity with running LLMs and agents in production (Amazon Bedrock or equivalent) and a strong interest in developing self-hosted serving on EKS/K8s.Advanced knowledge of Linux.Intermediate-advanced proficiency in English.
Nice to have:
AWS Certification (e.g., Machine Learning, DevOps Engineer, or Solutions Architect).Experience serving self-hosted models at scale on Kubernetes: distributed training and high-throughput, low-latency GPU inference.Experience optimizing LLM inference (e.g., vLLM, SGLang, TensorRT-LLM) and working with specialized AWS AI processors (Inferentia and Trainium).Experience with big data and streaming tools (Kafka, Spark) for real-time feature calculation.Experience in the energy or utilities sector, or familiarity with industry use cases (e.g., demand forecasting, predictive maintenance, fraud detection).
🚀 Do you recognize yourself in the "Who" and represent the "What"? Apply and let's talk!
We have 62,642 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 62,642 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume