Engineering Manager, KUBE Team

Cast Ai — Romania · Posted ~2 weeks ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Why Cast AI? Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale. By embedding autonomous decision-making directly into Kubernetes and cloud environments, Cast AI continuously optimizes performance, reliability, and efficiency in production. The old way doesn’t work. As Kubernetes and AI environments grow, manual decisions don’t. Cast AI replaces tickets, alerts, and manual tuning with continuous automation that adapts infrastructure as conditions change. Efficiency and cost savings follow naturally from that automation. Over 2,100 companies already rely on Cast AI, including Akamai, BMW, Cisco, FICO, HuggingFace, NielsenIQ, Swisscom, and TGS. Global team, diverse perspectives We’re headquartered in Miami, but our impact is international. We take a global and intentional approach to diversity. Today, Cast AI operates across 34 countries spanning Europe, North America, Latin America, and APAC, bringing a wide range of perspectives into how we build and lead. Unicorn momentum In January 2026, we achieved unicorn status with a strategic investment from Pacific Alliance Ventures, the corporate venture arm of Shinsegae Group (a $50+ billion Korean conglomerate). Our valuation now exceeds $1 billion, and we’re just getting started. Join us as we build the future of autonomous infrastructure. This is a location-specific opportunity. We are currently accepting applications from candidates residing in the following European countries: Bulgaria, Croatia, Estonia, Greece, Hungary, Latvia, Lithuania, Poland, Romania, Slovakia, Slovenia, and Ukraine. Why the KUBE team? KUBE is where Cast AI’s optimization stops being a recommendation and becomes real infrastructure. The Optimization Engine decides what to run and when; KUBE makes it happen — provisioning the compute, networking, and storage under every node, across EKS, GKE, AKS, and our own Cast AI Anywhere, and reconciling all of it continuously. The scale is the point. When demand spikes, KUBE brings up 15,000 nodes in a single cluster and gets every pod scheduled in under 45 minutes — and tears a cluster of that size back down in under 15 minutes — across multiple clouds at once, mixing instance types and spot capacity in a single operation. That is provisioning velocity the hyperscalers’ own tooling doesn’t deliver on its own, and KUBE owns the execution of every node in it. Across a fleet of thousands of customer clusters, the platform makes the same high-stakes Kubernetes decisions — add this node, drain that one, is this change safe — thousands of times a day, and this team is accountable for each one landing cleanly. The Problems Are Genuinely Hard, And Largely Undocumented We reverse-engineered how to attach independent, arbitrarily-typed nodes into managed clusters (EKS, GKE, AKS) that were never designed to allow it — and turned that into a reliable operation on every provider.We built a highly parallelized engine that manages and reconciles the diverse clusters of thousands of customers concurrently.We authored our own Terraform provider and a unified API that abstracts away the deep implementation differences between cloud providers.We own the duality of a Kubernetes node — every node is simultaneously a live cloud machine and a Kubernetes object, and KUBE keeps the two in lockstep through the full lifecycle: bootstrap, drain, delete, hibernate, resume.We are the SMEs for lower-level infrastructure — operating systems, cloud networking, storage, and virtualization — at Cast AI. Each sprint surfaces new hurdles at the intersection of Kubernetes internals and cloud-provider behavior. If you want ownership of a system operating at real scale, with room to shape both the architecture and the team building it, this is that role. Role overview We are looking for a hands-on Engineering Manager to join our Product team. It is a hybrid role: you will manage a team and spend at least half your time in the product code. Responsibilities as an Engineering Manager You own product development execution. Lead, mentor, and grow a team of talented software engineersCollaborate closely with the Product Owner and stakeholders to align on requirements and roadmapsEnsure quality, scalability, and resilience in all deliverablesFoster a culture of continuous improvement, innovation, and technical excellenceBe the role model and champion of Cast AI core values Responsibilities as a Software Engineer You will work in a highly skilled team across the full development lifecycle. Take on scaling challenges at fleet levelParticipate in feature brainstorming, requirement gathering, and solving customer needsDesign and architect software for scalability, maintainability, and performanceWrite clean, maintainable, documented code using best practicesEnsure reliability through test coverage and local testingContribute to CI/CD pipelines and the monitoring and alerting stackRespond to incidents to resolve customer issues or service disruptionsSupport the team’s code (on-call week every few months) Requirements Proven experience managing and leading a technical teamStrong software engineering skillsStrong verbal and written communication skillsIn-depth knowledge of Kubernetes and experience managing multi-cloud environments (AWS, GCP, Azure)Knowledge of cloud-native infrastructure, APIs, and cloud-provider-specific implementation differencesHands-on experience with IaC tools such as TerraformProficient in designing and managing CI/CD pipelines (GitLab CI, ArgoCD) with a GitOps approachStrong analytical, problem-solving, and troubleshooting skills What’s in it for you? Competitive salary (up to €10,000 gross, depending on the level of experience)Enjoy a flexible, remote-first global environment.Collaborate with a global team of cloud experts and innovators, passionate about pushing the boundaries of Kubernetes technology.Equity options.Get quick feedback with a fast-paced workflow. Most feature projects are completed in 1 to 4 weeks.Spend 10% of your work time on personal projects or self-improvement. Learning budget for professional and personal development – including access to international conferences and courses that elevate your skills.Annual hackathon to spark new ideas and strengthen team bonds.Team-building budget and company events to connect with your colleagues.Equipment budget to ensure you have everything you need.Extra days off to help maintain a healthy work-life balance. Hiring process Screening call with RecruiterHiring Manager interviewTechnical interview (system design)Live codingCulture Check interview with an executiveAs part of our standard hiring process, we would like to inform you that a background check may be conducted at the final stage of recruitment through our third-party provider, Checkr.Please note that Cast AI does not provide any form of visa sponsorship/work permit.