Technical Support Engineer II (Linux)

Vast Ai — United States · Posted ~1 hour ago

Mid Full-time

Skills

Linux Ubuntu Docker NVIDIA CUDA GPU Virtualization KVM Networking Hardware/BIOS/Firmware troubleshooting

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

We are looking for a Technical Support Engineer II (Linux) to join our growing team. You will be the go-to resource for resolving complex infrastructure issues, supporting our L1 team with escalated tickets. Your expertise in Linux, Docker, NVIDIA CUDA, GPU, virtualization, and networking will be key to ensuring our cloud infrastructure runs smoothly. We offer the chance to work on state-of-the-art AI systems alongside a globally distributed team, in a culture that values ownership and continuous learning.

Highlights

Work with cutting-edge AI systems, collaborate with a globally distributed team, contribute to democratizing AI computing, and join a company that values elegance, ownership, integrity, and continuous learning.

Description

About Us Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation. We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team. About The Role This is a technical support role focused on escalated infrastructure issues that go beyond frontline triage. You'll be the engineering resource our L1 support team leans on when tickets get complex: diagnosing and resolving issues across the full stack — hardware/BIOS/firmware, networking, Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM). You'll handle higher-complexity issues, own escalation resolution end-to-end, and contribute to internal documentation and runbooks. The best engineers in this role don't just resolve tickets — they build the tooling and runbooks that eliminate recurring ones. You'll collaborate directly with the engineering team and host support team on systemic issues. Strong technical depth and support experience are the primary requirements. You should be comfortable working autonomously across Ubuntu environments, diagnosing container and GPU issues, and communicating findings clearly to both technical and non-technical audiences. Vast.ai users or hosts strongly preferred. This role is full-time and onsite in our office in Westwood (LA) Schedule: Sunday - Thursday. Key Responsibilities Handle escalated support tickets, including GPU workload failures, container issues, networking problems, account infrastructure, and host-side configurationDiagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and virtualization environments (KVM)Troubleshoot network-layer issues: VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machinesInvestigate performance issues on GPU utilization, container resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecksAdvise suppliers (hosts) on installation best practices — hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performanceProvide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshootingWrite and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalationsBuild diagnostic and automation tooling in Python and Bash to reduce manual triage overheadCollaborate with the engineering team and infrastructure support team to flag and document systemic or recurring platform issuesAssist clients and infrastructure suppliers working with AI frameworks (TensorFlow, PyTorch) and GPU-accelerated workloadsProvide coverage for L1 support team overflow during peak periods or incidents, per a defined on-call rotation You Are Fluent in Linux — you navigate systems, read logs, and solve problems from the command line without hesitationMethodical and thorough: you gather data, dig into root causes, and don't settle for surface-level fixesA self-starter who can manage a queue of complex tickets with minimal supervisionAdaptable to a defined on-call rotation which may include weekend coverageA clear written communicator: able to explain technical findings and write useful internal documentationGenuinely curious about AI infrastructure, GPU computing, and distributed systems Must-Haves Solid Linux SysOps experience: Ubuntu Server, RHEL/CentOS, Debian; comfortable with systems, networking, storage, and permissionsProficiency with Docker: container debugging, Docker Compose, image management, cgroup resource limits, Docker storage/filesystem managementExperience with virtualization: Proxmox VE, VMware, or similar hypervisors; provisioning and troubleshooting VMsNetworking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules, and general L2/L3 troubleshootingHands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting (essential)Scripting in Python and Bash for automation and diagnostic toolingStrong English written communication: clear, professional, and technically preciseExperience providing technical support in a customer-facing or internal helpdesk contextAbility to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end Nice-to-Haves Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containersMonitoring and observability experience (Prometheus, Grafana)Relevant certifications: RHCSA, CompTIA Linux+, or similarKnowledge of the Vast.ai platform as a client or infrastructure supplier Annual Salary Range $90,000 – $150,000 + equity + benefits Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential. Benefits Comprehensive health, dental, vision, and life insurance401(k) with company matchMeaningful early-stage equityOnsite meals, snacks, and close collaboration with founders/tech leadersAmbitious, fast-paced startup culture where initiative is rewarded Compensation Range: $90K - $150K