Summary
Join a technically ambitious AI infrastructure team building the runtime layer behind high-performance enterprise AI applications. You will optimize large language model inference across distributed GPU systems, develop scalable model-serving services, and solve challenging problems spanning distributed systems, cloud infrastructure, and performance engineering.
Highlights
Work on highly technical AI infrastructure with a strong focus on performance, scalability, distributed systems, and efficient GPU utilization. The role offers exposure to production-scale AI workloads and modern cloud technologies.
Description
AI Runtime Engineer
Vienna, Austria โ Hybrid
AI Infrastructure | Distributed Systems | Large Language Models | High-Performance Computing
Our client, an innovative AI Infrastructure company based in Vienna, is building the runtime platform that powers high-performance AI applications used by enterprise customers across Europe.
They're looking for an AI Runtime Engineer to help optimise the systems responsible for serving Large Language Models across distributed GPU infrastructure, where every millisecond of latency and every percentage of GPU utilisation matters.
This is a deeply technical engineering role sitting at the intersection of distributed systems, cloud infrastructure, and modern AI.
Your Responsibilities
Develop high-performance runtime services responsible for serving production AI modelsOptimise inference performance across distributed GPU infrastructureBuild scalable model serving systems capable of handling enterprise AI workloadsDevelop platform services and performance tooling using Rust and PythonImprove scheduling, orchestration, and resource utilisation across Kubernetes clustersWork closely with AI Researchers and Machine Learning Engineers to optimise production deploymentProfile system performance and remove bottlenecks across networking, memory, and compute layersImprove platform observability, reliability, and operational efficiencyContribute to the architecture of the company's next-generation AI infrastructure platform
Experience Required
5+ years of experience in Backend Engineering, Platform Engineering, Distributed Systems, or AI InfrastructureStrong commercial experience developing software in Rust or modern systems programming languagesExcellent Python development skillsHands-on experience with Kubernetes in production environmentsExperience deploying or operating large-scale model serving infrastructureGood understanding of distributed systems, networking, concurrency, and cloud-native architecturesExperience working with GPU computing or performance-critical applicationsPassion for solving complex infrastructure and performance engineering challenges
Nice to Have
Experience with NVIDIA CUDA, Triton Inference Server, or vLLMExperience serving Large Language Models in productionKnowledge of distributed inference frameworksExperience with Ray, KServe, or Kubernetes GPU OperatorsFamiliarity with observability platforms such as Prometheus, Grafana, or OpenTelemetryPrevious experience in AI Infrastructure, HPC, or AI Developer Tools companies