Summary
An AI-focused engineering team is seeking a senior MLOps specialist to automate ML workflows, manage model lifecycles, and improve production reliability.
Highlights
Work on advanced machine learning infrastructure, automation, deployment reliability, and innovative AI solutions.
Description
What will you do?
ML Workflow Automation & Infrastructure
Design end-to-end architecture for the automated training of ML models Create data pipelines to build relevant datasets and data annotation flows Monitor ML model performance and data drift Handle versioning, deployment, and integration with the software team
Model Deployment & Lifecycle Management
Develop and manage CI/CD pipelines for building, testing, and deploying models Apply best practices for model versioning, rollback, and A/B testing to ensure reliable and accurate production releases
System Monitoring & Troubleshooting
Set up a robust monitoring system and develop automated alerting solutions to proactively identify issues in data pipelines, model training, validation, and data variation Promote MLOps best practices (Infrastructure as Code, reproducibility, security) and continuously improve internal processes to increase reliability and efficiency
Research & Innovation
Research and implement cutting-edge technologies to improve training efficiency (e.g., distributed training, HPC, multi-GPU strategies) for the research team Explore future MLOps frameworks and GPU-based cloud solutions as part of the scalability roadmap Collaborate with product managers and software engineers to identify AI/ML opportunities and develop innovative solutions
What are we looking for?
University degree, preferably in engineering (software, industrial, mechanical, process) or a related field Over 5 years of experience in MLOps or machine learning engineering, with a focus on deploying and managing deep learning models at scale Strong skills in Python, CI/CD pipelines, and ML frameworks (e.g., PyTorch, TensorFlow, OpenCV) for automating and scaling ML workflows Expertise in monitoring and alert automation for ML workflows, including data pipelines, training processes, and model performance (e.g., Prometheus, Grafana) Familiarity with distributed training techniques, multi-GPU strategies, and hardware optimization for deep learning Strong communication and interpersonal skills
What are we offering?
Meal tickets โ for your well-deserved breaks.
A place where your voice truly matters โ your ideas are heard and put into action.
Performance bonuses โ we value work done with passion.
A day off on your birthday โ celebrate it your way.
Private medical subscription โ for your health and your familyโs well-being.
Trainings and learning resources โ we invest in your continuous development.
Hybrid work model โ real balance between work and personal life.
Bookster subscription โ get your favorite books delivered for free.
A friendly, passionate, and solution-oriented team.
Opportunities to grow or change your role within the company.
Apply for this job!