Principal Software Engineer, Platform

Rafay Systems — United States · Posted ~8 hours ago

Lead

Skills

software architecture platform engineering Kubernetes multi-tenant SaaS multi-cloud architecture technical leadership hands-on software development SaaS multi-cloud GPU platforms virtualization

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Serve as a principal technical leader and hands-on architect for the backbone of a multi-tenant SaaS and Kubernetes-based platform operating across multiple clouds. Shape technical direction, design critical platform components, and contribute directly to advanced platform engineering involving container orchestration, GPU infrastructure, and virtualization.

Highlights

Principal-level technical leadership with hands-on architecture and development of the backbone of advanced multi-cloud platforms, plus opportunities for innovation and career advancement.

Description

Principal Software Engineer (Platform) We are looking for a Principal Software Engineer who can make significant contributions to the design and development of the backbone of our multi-tenant SaaS, GPU PaaS, and Virtualization Kubernetes Operations Platform for a multi-cloud environment. Rafay is at the forefront of Kubernetes technology and we offer unique opportunities to develop new technology and be part of a team that encourages positive change through outside-of-the-box thinking. We hold high expectations for ourselves and challenge team members to continually seek improvement. Rafay offers opportunities to work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. As the first of our kind, we are truly in a class of our own. As a Principal Engineer, you will serve as a technical leader and hands-on architect, setting the technical direction for some of the most critical services on the platform and multiplying the impact of the engineering teams around you. Responsibilities Design and Implement core architectural components for some of the most critical platform services of a multi-tenant distributed SaaS, GPU PaaS, and Virtualization platformSet the technical vision and long-term architectural direction for major platform areas, and drive alignment across engineering teamsBuild highly modular and scalable components and services for the platformPerform R&D, feasibility analysis on latest technologies and newer versions of frameworks and libraries on an ongoing basisAssist operations and solutions teams with deployment and stability of production systemsCollaborate with other team members and stakeholders including product management, UI designers and QAParticipate in and lead code reviews and design reviews, and establish engineering best practices and standards across teamsMentor and provide technical guidance to senior and junior engineers, raising the technical bar across the organizationEnable a customer obsessed environment where team members can relentlessly champion and advocate for our customers in representing their issues to engineering teams and be a change agent to develop innovative ways to resolve their issuesEngineer and implement control plane services for GPU task scheduling, intelligent placement, and lifecycle handling across multi-cloud, diverse hardware infrastructureBuild highly scalable inference serving systems that deliver strict latency and throughput targets for multi-tenant deployments while optimizing GPU usage, capacity planning, and auto-scalingArchitect solution capabilities for model registries, artifact delivery, versioning strategies, blue-green or canary deployments, and concurrent LoRA inference execution across diverse GPU infrastructure Skills and Qualifications 12+ years of experience in building and delivering large enterprise applications for customersDeep understanding of distributed systems fundamentals, high availability and scalability principlesExpert knowledge of one or more of the following programming languages Golang, PythonExperience developing Micro-servicesStrong troubleshooting and debugging skillsStrong understanding of multi-tenant isolation: per-tenant quotas, noisy-neighbor mitigation, and data, workload, and network isolation boundariesHands-on experience developing services on a public cloud platform (e.g., AWS, Azure, GCP)Practical knowledge of networking protocols (TCP/IP, HTTP) and standard network architecturesExperience with containers and orchestration technologies like KubernetesExperience with GPU infrastructure, accelerated computing, or GPU scheduling/orchestration (e.g., NVIDIA GPU Operator, MIG, multi-tenant GPU sharing)Experience with model registry and multi-cluster artifact distribution: model versioning, promotion across environments, and efficient replication of large artifacts to remote clusters is a plusUnderstanding of LLM inference performance characteristics: continuous batching, KV cache management, prefill vs. decode behavior, quantization, and tensor/pipeline parallelism is a plusUnderstanding of topology-aware scheduling and high-performance interconnects: NVLink, NCCL, RDMA/InfiniBand, GPUDirect, and their impact on placement decisions is a plusExperience working and delivering features/enhancements/critical fixes for customer found issuesDemonstrated ability to influence technical direction and drive consensus across multiple teams and stakeholders WHY JOIN RAFAY Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.