Summary
✨ AI‑Generated
A full-time hybrid DevOps / Platform Engineer role focused on owning production Kubernetes infrastructure and multi-cloud environments. You’ll build infrastructure as code, maintain CI/CD pipelines, deploy containerized services, implement comprehensive observability, and support demanding AI/ML workloads including model serving and GPU processing. The role emphasizes secure, reliable, scalable, and cost-efficient production systems.
Highlights
Own and evolve cloud and Kubernetes infrastructure for production AI workloads, with strong focus on scalability, security, reliability, observability, and cost efficiency. Work closely with product, engineering, and ML teams in a hands-on platform role.
Description
About MarvelX
MarvelX builds AI agents for regulated industries such as insurance and financial services.
Our systems run in real production environments where reliability, security, and observability matter as much as model quality.
We’re looking for a DevOps / Platform Engineer to own and evolve our cloud and Kubernetes infrastructure.
You will work closely with Product, Engineering, and ML to keep our platform scalable, secure, and production-ready as we grow.
This is a full-time, hybrid role (office-first, up to 1 day/week remote).
What you will do
Own and operate Kubernetes clusters running production workloadsBuild and maintain infrastructure across AWS, GCP, and AzureDefine infrastructure as code using TerraformBuild and maintain GitLab CI/CD pipelinesContainerize and deploy services using DockerImplement observability: metrics, logs, tracing, alertingSupport AI/ML workloads (model serving, batch jobs, GPU workloads)Improve reliability, security posture, and cost efficiencyPartner with engineering to design production-ready architectures
What we’re looking for
Strong hands-on experience operating Kubernetes in productionExperience with at least one major cloud provider (AWS, GCP, or Azure); multi-cloud experience is a plusSolid Terraform experience managing real infrastructureGitLab CI/CD or similar pipeline tooling experienceStrong Docker fundamentalsExperience with observability tooling (Prometheus, Grafana, ELK, OpenTelemetry, or similar)Linux and scripting skills (Bash, Python, etc.)Interest or experience supporting AI/ML systems
Nice to have
GPU scheduling / ML platform experienceHelm or KustomizeExposure to security and compliance environments (SOC 2, ISO 27001)
Why join
Own the platform of a production AI companySmall, senior, highly technical teamDirect impact on reliability, security, and scaleCompetitive salary + equity