Senior SW Engineer – AI Infrastructure & Optimization

Dou Eu — Poland · Posted ~22 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Working mode: 📍Hybrid in Krakow (2 days/week office) We are looking for a Senior Software Engineer to help build and optimize large-scale, high-performance GenAI infrastructure and inference systems on Kubernetes. As AI workloads increasingly move toward Kubernetes-native infrastructure, we are building systems that support distributed inference, performance optimization, reliability, observability, and production-grade deployment at scale. This role is ideal for an engineer who can reason deeply about systems, performance, tradeoffs, and reliability, and who is comfortable owning difficult technical decisions end-to-end. You will work across inference serving, distributed systems, optimization, and Kubernetes-native AI infrastructure. What You’ll Do Build and optimize high-performance Kubernetes-native GenAI inference systemsWork with modern inference stacks such as vLLM, SGLang, TensorRT-LLM, and related toolingWork with Kubernetes-native distributed LLM inference frameworks such as llm-d and NVIDIA DynamoDesign and implement optimization algorithms and performance improvementsImprove reliability, observability, deployment, and operational maturity of AI systemsMake architectural decisions and take ownership of technical outcomesCollaborate with a small, senior engineering team focused on performance and production quality Required Qualifications Minimum 5 years of experience as a Software Engineer, with strong software engineering and system design skills.Programming experience in Go and PythonHands-on experience with the Kubernetes ecosystem, including Operators, service meshes, GitOps, Gateway API, and OpenTelemetryExperience with cloud platformsStrong understanding of optimization algorithms and performance engineeringAbility to independently drive technical initiatives from concept to productionStrong systems thinking and debugging skillsComfort operating in environments with high autonomy and responsibility Nice to Have Experience with modern LLM inference frameworks such as vLLM, SGLang, or TensorRT-LLMExperience with distributed LLM inference frameworks such as llm-d or NVIDIA DynamoContributions to open-source Kubernetes or ML infrastructure projectsGPU performance optimization and profiling experienceFamiliarity with CUDA, NCCL, or Triton kernelsExperience running GenAI systems at scale in production