Platform Engineer – AI/ML Infrastructure

Testqtech — Sweden · Posted ~5 hours ago

Senior Full-time

Skills

Kubernetes Docker Terraform OpenTofu Pulumi GitLab CI GitHub Actions ArgoCD cloud infrastructure GPU provisioning cluster autoscaling AI/ML workloads Kubeflow MLflow AWS Azure GCP Pinecone Milvus Qdrant

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Own platform infrastructure for demanding AI and machine-learning workloads. You will manage containerized services, automate environments with infrastructure-as-code, build CI/CD and GitOps workflows, provision cloud GPUs, operate autoscaling clusters, and support vector and distributed storage systems.

Highlights

Build and operate modern AI/ML infrastructure using Kubernetes, infrastructure-as-code, CI/CD, cloud platforms, GPU resources, MLOps tooling, and vector databases.

Description

Required Skills & Qualifications Experience: 4+ years in platform engineering, DevOps, or infrastructure engineering, with at least 1-2 years supporting AI/ML workloads.Containerization & Orchestration: Advanced knowledge of Kubernetes (AKS, EKS, or GKE) and Docker for managing containerized AI services.Infrastructure as Code (IaC): Deep expertise with tools like Terraform, OpenTofu, or Pulumi to automate environment setup.CI/CD & MLOps Tools: Hands-on experience with GitLab CI, GitHub Actions, ArgoCD, and specialized platforms like Kubeflow or MLflow.Cloud & GPU Provisioning: Experience managing cloud services (AWS, Azure, or GCP) with an emphasis on provisioning GPU instances and handling cluster autoscaling.Data & Storage Systems: Experience setting up or maintaining vector databases (e.g., Pinecone, Milvus, Qdrant) and distributed storage solutions.