AI Runtime Engineer

Xpertdirect โ€” Austria ยท Posted ~1 day ago

Hybrid

Skills

Distributed systems Cloud infrastructure Large language models GPU infrastructure AI model serving Inference optimization High-performance computing Large Language Models GPU Model serving

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Join a technically ambitious AI infrastructure team building the runtime layer behind high-performance enterprise AI applications. You will optimize large language model inference across distributed GPU systems, develop scalable model-serving services, and solve challenging problems spanning distributed systems, cloud infrastructure, and performance engineering.

Highlights

Work on highly technical AI infrastructure with a strong focus on performance, scalability, distributed systems, and efficient GPU utilization. The role offers exposure to production-scale AI workloads and modern cloud technologies.

Description

AI Runtime Engineer Vienna, Austria โ€” Hybrid AI Infrastructure | Distributed Systems | Large Language Models | High-Performance Computing Our client, an innovative AI Infrastructure company based in Vienna, is building the runtime platform that powers high-performance AI applications used by enterprise customers across Europe. They're looking for an AI Runtime Engineer to help optimise the systems responsible for serving Large Language Models across distributed GPU infrastructure, where every millisecond of latency and every percentage of GPU utilisation matters. This is a deeply technical engineering role sitting at the intersection of distributed systems, cloud infrastructure, and modern AI. Your Responsibilities Develop high-performance runtime services responsible for serving production AI modelsOptimise inference performance across distributed GPU infrastructureBuild scalable model serving systems capable of handling enterprise AI workloadsDevelop platform services and performance tooling using Rust and PythonImprove scheduling, orchestration, and resource utilisation across Kubernetes clustersWork closely with AI Researchers and Machine Learning Engineers to optimise production deploymentProfile system performance and remove bottlenecks across networking, memory, and compute layersImprove platform observability, reliability, and operational efficiencyContribute to the architecture of the company's next-generation AI infrastructure platform Experience Required 5+ years of experience in Backend Engineering, Platform Engineering, Distributed Systems, or AI InfrastructureStrong commercial experience developing software in Rust or modern systems programming languagesExcellent Python development skillsHands-on experience with Kubernetes in production environmentsExperience deploying or operating large-scale model serving infrastructureGood understanding of distributed systems, networking, concurrency, and cloud-native architecturesExperience working with GPU computing or performance-critical applicationsPassion for solving complex infrastructure and performance engineering challenges Nice to Have Experience with NVIDIA CUDA, Triton Inference Server, or vLLMExperience serving Large Language Models in productionKnowledge of distributed inference frameworksExperience with Ray, KServe, or Kubernetes GPU OperatorsFamiliarity with observability platforms such as Prometheus, Grafana, or OpenTelemetryPrevious experience in AI Infrastructure, HPC, or AI Developer Tools companies