Summary
✨ AI‑Generated
A highly hands-on senior DevOps engineering role for building and operating production-grade machine learning and data platforms. Responsibilities include managing Kubernetes clusters, scaling infrastructure with Infrastructure as Code, and engineering CI/CD pipelines for machine learning workflows.
Highlights
Highly hands-on senior DevOps role focused on AWS, Kubernetes, Infrastructure as Code, production operations, and MLOps automation for production-grade machine learning and data platforms.
Description
Senior AWS DevOps Engineer (100% Hands-on)
We are seeking a high-octane, 100% hands-on AWS DevOps Engineer to join our engineering team.
This is not an architectural or high-level strategy role; we need a "builder" who thrives in the terminal, writes clean Infrastructure as Code, and can manage the end-to-end lifecycle of production-grade ML and data platforms.
The ideal candidate has deep expertise in Kubernetes (EKS) and a proven track record of automating complex MLOps workflows.
AWS Certifications are a significant plus.
Core Responsibilities (The "Must-Haves")
EKS Mastery: Lead the end-to-end implementation and day-to-day operations of Amazon EKS, including cluster provisioning, networking, and hosting complex microservices/apps.Infrastructure as Code: Maintain and scale our environment using Terraform as the primary provider.MLOps Pipeline Engineering: Build and manage CI/CD pipelines specifically for Machine Learning, including automated model registry updates, feature store synchronization, and automated deployments.Cloud Security & Identity: Implement rigorous IAM best practices specifically for ML workloads and ensure data isolation through VPCs, subnets, and private endpoints.Data Persistence & Governance: Manage S3 (lifecycle policies, performance tiering), Secrets Manager, KMS encryption, and Parameter Store.Observability: Own the monitoring stack using CloudWatch, Prometheus, and Grafana to ensure system health and proactive alerting.
Technical Requirements & Skills
Infrastructure & Platforms:
Primary: Amazon EKS, Terraform, Airflow (Orchestration).Secondary: Kubeflow or equivalent MLOps tooling.
Database & Analytics
Relational: RDS/Aurora (PostgreSQL/MySQL).NoSQL: DynamoDB (expert-level capacity planning and data modeling).Data Engineering: AWS Glue (ETL), EMR (EMR on EKS/Serverless).
Machine Learning Operations
Deep experience with Amazon SageMaker (training jobs, endpoints, pipelines, and notebooks).
Engineering Principles
We expect you to live and breathe the AWS 5 Pillars (The Well-Architected Framework):
Cost Optimization: Proactive management of cloud spend.
Performance Efficiency: Selection of evolved resource types based on workload.
Reliability: Designing for recovery and change management.
Resiliency: Implementing Multi-AZ/Multi-Region patterns.
Security: Operating under the principle of Least Privilege and cloud security best practices.
Qualifications
Minimum 5+ years of experience in an AWS DevOps or Site Reliability role.Proven experience supporting Data Science teams and ML production environments.Ability to troubleshoot complex networking and container orchestration issues in real-time.Must be 100% hands-on.