Summary
✨ AI‑Generated
A technology company is seeking a backend engineer to build reliable services powering advanced AI workloads. The role focuses on designing scalable APIs, improving inference performance, optimizing infrastructure components, and creating production-ready systems for machine learning teams.
Highlights
Opportunity to build high-performance AI infrastructure, optimize modern machine learning workloads, and work on scalable backend systems with significant performance impact.
Description
Company Description Throttle helps companies running large language models on their own infrastructure dramatically reduce inference costs, often by 40–60%.
Its drop-in semantic cache removes redundant compute on repeat queries, delivering savings without requiring changes to existing code.
Throttle is designed to integrate seamlessly with popular inference frameworks such as vLLM, SGLang, Ollama, and LMDeploy.
Early benchmarks show 2–3x throughput gains on real-world workloads, making Throttle an attractive solution for teams optimizing performance and cost.
The company is focused on building robust, production-ready tooling for AI infrastructure teams.
Role Description As a Backend Engineer, AI Inference at Throttle, you will design, build, and maintain backend services that power semantic caching for large language model inference.
You will develop high-performance APIs and infrastructure components, optimize inference pipelines, and ensure reliability, scalability, and observability across production workloads.
Day-to-day, you will collaborate closely with AI researchers and infrastructure engineers to integrate with frameworks like vLLM, SGLang, Ollama, and LMDeploy, and to implement caching, persistence, and routing logic.
You will also participate in code reviews, improve developer tooling, and help shape best practices for backend systems supporting AI workloads.
This is a full-time, remote, unpaid role based in the San Francisco Bay Area.
Qualifications
Strong Back-End Web Development and Software Development skills, with experience building scalable, production-grade services.Proficiency in Object-Oriented Programming (OOP) and general Programming, using languages commonly employed in backend and AI infrastructure (e.g., Python, Go, Rust, or similar).Understanding of Front-End Development fundamentals sufficient to collaborate with product and frontend teams on APIs and integration points.Experience with distributed systems, performance optimization, and high-throughput services, ideally in AI or data-intensive environments.Familiarity with LLM inference frameworks (such as vLLM, SGLang, Ollama, LMDeploy) or similar ML serving platforms.Comfort with cloud infrastructure, containers, and orchestration tools (e.g., Kubernetes, Docker) and modern CI/CD practices.Strong problem-solving skills, clear communication, and the ability to work collaboratively in an on-site, fast-paced startup environment.Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.