Staff Software Engineer - AI Compute, Together Cloud

Talenthopllc โ€” United States ยท Posted ~3 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

This is a Fully Remote Job About Our Client: The organization operates in the AI cloud infrastructure industry, focused on addressing the challenge of providing high-performance, AI-ready GPU clusters and advanced virtualization of machine learning hardware. It develops an end-to-end platform that supports the generative AI lifecycle, including inference, reinforcement learning, and fine-tuning products. The company delivers these capabilities through a self-serve cloud console and manages infrastructure across multiple data centers and hundreds of thousands of GPUs. About the Opportunity: The STAFF SOFTWARE ENGINEER - AI COMPUTE is responsible for setting the technical direction and building major components of the next-generation AI cloud platform. This role impacts the organization by architecting and developing a highly available, global cloud infrastructure with advanced virtualization of ML hardware. The engineer will lead design and development efforts from hardware bring-up to global management, ensuring the platform's reliability, scalability, and performance. This position is critical for enabling the organization's SaaS products and self-serve offerings and for advancing the overall AI cloud platform. Responsibilities: Own GPU and network virtualization stack, ensuring high performance, portability, and strong isolation across diverse hardware.Architect and develop in-data-center infrastructure-as-a-service (IaaS) layers, including compute, storage, and network provisioning.Design GPU scheduling and global management planes to support on-demand and reserved clusters across data centers.Architect monitoring and automated fault remediation systems for maintaining service availability.Lead cross-team technical direction, design reviews, and integration efforts.Mentor engineers and contribute to team growth and expertise in virtualization and GPU infrastructure.Develop testing frameworks, tools, and documentation to enhance system robustness and usability. Requirements: 7+ years of professional software development experience with expert proficiency in backend languages (Golang preferred).Proven experience owning architecture and production deployment of large distributed systems at scale.Deep knowledge of globally distributed, high-performance microservice architectures on cloud platforms.Expert understanding of compute, networking, and storage systems including concurrency and memory management.Demonstrated technical leadership and ability to align multiple teams.Strong communication and diplomacy skills for technical and non-technical collaboration.Experience with infrastructure automation, observability, and continuous integration/deployment systems. Pay Range and Compensation Package: The pay range and compensation package for this role will be determined based on the candidate's experience, skills, and other relevant factors. Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin. Note: TalentHop is a recruitment partner of this role. Please note that all employment decisions, including candidate assessment, interviews, hiring, compensation, and employment terms, are made exclusively by the hiring employer.