Senior/Staff Machine Learning Engineer

Torentify Official — United States · Posted ~3 hours ago

Senior Remote

Skills

Distributed systems Large-scale machine learning training Distributed ML training Model parallelism Machine learning engineering Machine Learning Distributed ML Foundation models

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a research-driven team developing new infrastructure for distributed training of frontier machine learning models. You’ll work on distributed systems, large-scale ML training, and model parallelism, helping create technology capable of training sophisticated models across consumer-grade internet connections. This is an opportunity for an experienced ML engineer to tackle deep technical challenges at the intersection of distributed computing and AI research.

Highlights

Senior/Staff opportunity focused on cutting-edge distributed machine learning, foundation-model training, and novel approaches to large-scale model development over distributed networks.

Description

## About the Company Pluralis Research conducts foundational research into Protocol Learning, a distributed approach to training foundation models in which no single participant has or can obtain a complete copy of the model. The company is focused on enabling community-trained and community-owned frontier models with self-sustaining economics. Pluralis Research is backed by Union Square Ventures and other tier-1 investors and brings together ML researchers and engineers with experience from leading technology companies and startups. ## About the Role Pluralis Research is seeking **Senior/Staff Machine Learning Engineers** with **5+ years of experience in distributed systems and large-scale machine learning training**. In this role, you will help build a novel technical foundation for training distributed ML models over consumer-grade internet connections. The role is suited to engineers with deep expertise in **distributed systems, distributed ML training, model parallelism, networking, GPU optimization, and production Python**. You will design systems capable of operating across heterogeneous hardware and unreliable networks while maintaining performance, resilience, and efficient communication between participants. ## Key Responsibilities ### Distributed Training Architecture & Optimization * Design and implement large-scale distributed training systems for heterogeneous hardware operating under low-bandwidth and high-latency network conditions. * Develop and optimize model-parallel training strategies, including data, tensor, and pipeline parallelism. * Implement custom sharding techniques designed to minimize communication overhead. * Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes. * Build robust checkpointing, state synchronization, and recovery mechanisms for long-running and fault-prone training jobs. * Develop monitoring and metrics systems to track training progress, model quality, and system bottlenecks. ### Decentralized Networking & Resilience * Architect resilient distributed training systems that can operate despite node failures and network partitions. * Support systems where participants can dynamically join or leave the network. * Design and optimize peer-to-peer topologies for decentralized coordination across non-co-located nodes. * Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management. * Profile and optimize communication patterns to reduce latency and bandwidth usage across multi-participant environments. ## Required Qualifications * **5+ years of experience in distributed systems and large-scale machine learning training.** * Strong experience building and operating distributed systems in production. * Hands-on experience with distributed training frameworks such as **FSDP, DeepSpeed, Megatron, or similar**. * Deep understanding of **model parallelism**, including data, tensor, and pipeline parallelism. * Expert-level **Python** experience in production environments. * Experience with Python concurrency, error handling, retry logic, and clean software architecture. * Strong networking fundamentals, including **P2P systems, gRPC, routing, NAT traversal, and distributed coordination**. * Experience optimizing GPU workloads, memory management, and large-scale compute efficiency. ## Preferred Qualifications The original job description does not specify separate preferred or nice-to-have qualifications. All qualifications listed above are presented as requirements in the source description. ## Skills & Competencies * Distributed Systems * Distributed Machine Learning * Large-Scale ML Training * Machine Learning Engineering * Model Parallelism * Data Parallelism * Tensor Parallelism * Pipeline Parallelism * FSDP * DeepSpeed * Megatron * Python * GPU Optimization * GPU Utilization * Memory Management * Large-Scale Computing * Custom Model Sharding * Checkpointing * State Synchronization * Fault Recovery * Monitoring and Metrics * Peer-to-Peer (P2P) Networking * gRPC * Routing * NAT Traversal * Peer Discovery * Dynamic Routing * Distributed Coordination * Connection Lifecycle Management * Network Optimization * Latency and Bandwidth Optimization * Production Systems * Resilient Infrastructure ## Education & Experience **Education:** No specific education or degree requirement is stated in the original job description. **Experience:** * 5+ years of experience in distributed systems and large-scale ML training. * Production experience building and operating distributed systems. * Hands-on experience with distributed ML training frameworks and model-parallel training. * Production-level Python experience. * Experience with distributed networking and GPU workload optimization. ## Work Arrangement & Schedule * **Location:** California, Missouri, United States * **Work Arrangement:** Remote-first * **Employment Type:** Senior/Staff engineering role * **Remote Work:** Remote-first * **Optional Office Access:** Melbourne hub * No specific weekly hours or weekend requirements are stated in the original job description. **Location note:** The source listing identifies the job location as **California, MO, US**, while the company description states that the role is remote-first with optional access to a **Melbourne hub**. The source also mentions visa sponsorship for exceptional candidates and a competitive base salary for senior engineering roles in Australia. These location-related details appear inconsistent, so the listed job location and remote-first arrangement have been preserved without inventing a correction. ## Compensation & Benefits The original job description does not provide a specific salary range. It states that the company offers: * Equity-heavy compensation with meaningful ownership in a mission-driven company * Competitive base salary for senior engineering roles in Australia * Visa sponsorship available for exceptional candidates * Remote-first work with optional access to the Melbourne hub * Opportunity to work with a technical team whose members have experience at Google, Amazon, Microsoft, and leading startups ## Compliance / Additional Information Pluralis Research describes Protocol Learning as an approach intended to support community-trained and community-owned frontier models and reduce concentration of model development, access, and economic value among a small number of large corporations. The company is backed by **Union Square Ventures and other tier-1 investors**.