Summary
✨ AI‑Generated
A software engineering role focused on building production-grade machine learning infrastructure, inference systems, optimization pipelines, and tools for advanced AI workloads.
Highlights
Work on cutting-edge ML infrastructure connecting research models with production systems and hardware optimization.
Description
At Orient Path, we partner with high-growth technology companies globally.
One of our clients is an early-stage deep-tech AI company developing a fundamentally new approach to LLM model compression — making foundation models smaller, faster, and more efficient to run in production.
Their research team develops compression methods.
Now they are building the engineering layer that turns this research into reliable, high-performance ML infrastructure.
What are we building?
Production infrastructure around LLM compression and inference.
The engineering work sits between models and hardware: quantization, pruning and retraining pipelines, inference frameworks, GPU execution, performance profiling, benchmarking, and production tooling.
Who are we looking for?
A Software Engineer — ML Systems & AI Infrastructure who has already built and run the systems behind modern ML/LLM workloads.
This is not a general backend, MLOps, or AI application role.
You should be strong in at least one of these areas:
Training: distributed model training, FSDP/ZeRO/DDP, mixed precision, checkpointing and multi-GPU workloadsML Infrastructure: LLM serving, GPU clusters, distributed execution, profiling, throughput and production ML systemsGPU Performance / Kernels: GPU profiling, memory and compute optimization, kernel performance, CUDA/Triton or similar low-level workAcross all three, we expect strong software engineering fundamentals and hands-on ownership of real ML systems.
Tech stack:
Python, PyTorch, vLLM, Hugging Face, TensorRT-LLM, llama.cpp, CUDA/Triton, GPU profiling & benchmarking, distributed GPU systems, CI/CD
What we expect from candidates:
Strong production Python and software engineering fundamentalsHands-on experience building or optimizing ML/LLM systems, not simply using AI APIs or modelsStrong understanding of GPU performance: memory, compute, profiling, bottlenecks and hardware constraintsExperience with LLM inference, distributed training, GPU infrastructure, or kernel optimizationExperience with PyTorch and modern ML/LLM frameworksAbility to investigate performance problems hands-on: profile → identify the bottleneck → implement the change → measure the resultExperience turning research/experimental code into reusable, reliable engineering systemsUnderstanding of quantization, pruning, mixed precision, model compression or other model-efficiency techniques
Especially relevant experience:
CUDA/Triton or GPU kernels; vLLM/TensorRT-LLM/llama.cpp; FSDP/ZeRO/DDP; NCCL and multi-GPU systems; profiling with Nsight or similar tools; open-source ML infrastructure contributions.
What we offer:
€70K–120K base salary, depending on experienceEquity participationVienna / remote within EuropeVisa sponsorship & relocation supportDirect work with founders and researchersHigh ownership in a small, deeply technical team
If this sounds close to what you've actually been building, I'd love to talk.