DevOps Specialist

Hermes Corporate — Italy · Posted ~22 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Senior DevOps Engineer – AI & HPC Infrastructure We are looking for a Senior DevOps Engineer to lead the design, operation, and evolution of a cutting-edge AI and High-Performance Computing (HPC) infrastructure. This is not a traditional DevOps role. We're looking for someone who has built and operated mission-critical infrastructure at scale, thrives in complex environments, and can take full ownership of platform reliability, automation, and performance. You will work at the intersection of AI, cloud infrastructure, HPC, and MLOps, enabling world-class engineering and machine learning teams. What You'll Do Own and evolve production-grade AI and HPC infrastructure.Design, build, and optimize highly automated CI/CD and deployment workflows.Improve platform reliability, scalability, observability, and operational excellence.Build automation and internal tooling using Python and Bash.Manage complex Linux-based environments, ensuring performance and stability.Operate distributed and containerized workloads across cloud infrastructure.Partner closely with AI/ML teams to enable efficient model development and deployment.Define infrastructure best practices and drive continuous platform improvements. What We're Looking For Required Qualifications Extensive experience in Senior DevOps, Platform Engineering, or Infrastructure Engineering roles.Demonstrated experience operating and supporting High-Performance Computing (HPC) environments in production.Deep expertise in Linux administration, troubleshooting, security, and performance tuning.Advanced proficiency in Python and Bash for automation, tooling, and operational workflows.Strong experience designing, implementing, and maintaining modern CI/CD pipelines, including GitHub Actions.Extensive experience with AWS in complex, multi-environment production architectures.Proven experience managing containerized and distributed systems.Track record of owning infrastructure end-to-end and driving operational excellence.Excellent troubleshooting, analytical, and problem-solving skills with a production-first mindset. Preferred Qualifications Experience managing GPU infrastructure, including NVIDIA drivers and CUDA.Strong understanding of MLOps, including model deployment, lifecycle management, and monitoring.Experience supporting AI/ML engineering teams and bridging research with production infrastructure.Hands-on experience with Infrastructure as Code (IaC) and large-scale infrastructure automation. Ideal Candidate We're looking for a highly experienced engineer who has worked in demanding production environments and is comfortable taking ownership of critical infrastructure. The ideal candidate combines deep technical expertise with a pragmatic approach to engineering, automation, and operational excellence, and enjoys solving complex infrastructure challenges that support advanced AI workloads.