Summary
✨ AI‑Generated
A founding engineering role focused on designing and optimizing infrastructure for large-scale AI workloads. The position involves architecture, performance tuning, and building high-impact machine learning systems.
Highlights
Foundational engineering role building advanced AI infrastructure with significant technical ownership and impact.
Description
Company Description Throttle helps companies significantly reduce the cost of running large language models on their own infrastructure, often cutting inference spend by 40–60%.
The platform provides a drop-in semantic cache that eliminates redundant compute on repeat queries without requiring changes to existing code.
Throttle is compatible with leading LLM serving frameworks such as vLLM, SGLang, Ollama, and LMDeploy.
Early benchmarks demonstrate 2–3x throughput gains on real-world workloads, making Throttle a compelling solution for teams scaling AI systems.
The company is focused on building high-leverage infrastructure for developers who care about performance, reliability, and efficiency.
Role Description As a Founding ML Systems Engineer at Throttle, you will design, build, and optimize the core infrastructure that powers our semantic caching platform for LLM workloads.
You will work on end-to-end systems engineering tasks, including architecture design, performance tuning, resource management, and integration with popular serving frameworks.
Day-to-day work will involve implementing robust caching strategies, profiling and troubleshooting production systems, improving reliability and observability, and collaborating closely with founders and customers to translate real-world requirements into scalable solutions.
This is a full-time, unpaid, remote work where you will help define engineering best practices, contribute to technical direction, and shape the foundation of Throttle’s engineering culture.
Qualifications
Strong systems engineering skills, including experience with distributed systems, performance optimization, and resource management.Hands-on expertise in systems design, covering scalable architectures, fault tolerance, and high-throughput data processing.Experience with system administration tasks such as monitoring, configuration management, and deployment in Linux-based environments.Ability to provide technical support and troubleshooting for production systems, including diagnosing bottlenecks and resolving incidents.Proficiency with at least one systems-level programming language (e.g., Rust, C++, Go) and familiarity with Python for ML-related tooling.Experience working with LLM or ML inference infrastructures, model serving frameworks (e.g., vLLM, SGLang, Ollama, LMDeploy), or similar platforms.Comfort with observability tools (logging, metrics, tracing) and cloud or on-prem infrastructure (e.g., Kubernetes, containerization, GPU scheduling).Strong collaboration and communication skills, with the ability to work closely with a small founding team and early customers.Bachelor’s degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.Startup experience or interest in early-stage environments, with a willingness to take ownership and iterate quickly.