Summary
✨ AI‑Generated
Join a research-driven team developing new infrastructure for distributed training of frontier machine learning models. You’ll work on distributed systems, large-scale ML training, and model parallelism, helping create technology capable of training sophisticated models across consumer-grade internet connections. This is an opportunity for an experienced ML engineer to tackle deep technical challenges at the intersection of distributed computing and AI research.
Highlights
Senior/Staff opportunity focused on cutting-edge distributed machine learning, foundation-model training, and novel approaches to large-scale model development over distributed networks.
Description
## About the Company
Pluralis Research conducts foundational research into Protocol Learning, a distributed approach to training foundation models in which no single participant has or can obtain a complete copy of the model.
The company is focused on enabling community-trained and community-owned frontier models with self-sustaining economics.
Pluralis Research is backed by Union Square Ventures and other tier-1 investors and brings together ML researchers and engineers with experience from leading technology companies and startups.
## About the Role
Pluralis Research is seeking **Senior/Staff Machine Learning Engineers** with **5+ years of experience in distributed systems and large-scale machine learning training**.
In this role, you will help build a novel technical foundation for training distributed ML models over consumer-grade internet connections.
The role is suited to engineers with deep expertise in **distributed systems, distributed ML training, model parallelism, networking, GPU optimization, and production Python**.
You will design systems capable of operating across heterogeneous hardware and unreliable networks while maintaining performance, resilience, and efficient communication between participants.
## Key Responsibilities
### Distributed Training Architecture & Optimization
* Design and implement large-scale distributed training systems for heterogeneous hardware operating under low-bandwidth and high-latency network conditions.
* Develop and optimize model-parallel training strategies, including data, tensor, and pipeline parallelism.
* Implement custom sharding techniques designed to minimize communication overhead.
* Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.
* Build robust checkpointing, state synchronization, and recovery mechanisms for long-running and fault-prone training jobs.
* Develop monitoring and metrics systems to track training progress, model quality, and system bottlenecks.
### Decentralized Networking & Resilience
* Architect resilient distributed training systems that can operate despite node failures and network partitions.
* Support systems where participants can dynamically join or leave the network.
* Design and optimize peer-to-peer topologies for decentralized coordination across non-co-located nodes.
* Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.
* Profile and optimize communication patterns to reduce latency and bandwidth usage across multi-participant environments.
## Required Qualifications
* **5+ years of experience in distributed systems and large-scale machine learning training.**
* Strong experience building and operating distributed systems in production.
* Hands-on experience with distributed training frameworks such as **FSDP, DeepSpeed, Megatron, or similar**.
* Deep understanding of **model parallelism**, including data, tensor, and pipeline parallelism.
* Expert-level **Python** experience in production environments.
* Experience with Python concurrency, error handling, retry logic, and clean software architecture.
* Strong networking fundamentals, including **P2P systems, gRPC, routing, NAT traversal, and distributed coordination**.
* Experience optimizing GPU workloads, memory management, and large-scale compute efficiency.
## Preferred Qualifications
The original job description does not specify separate preferred or nice-to-have qualifications.
All qualifications listed above are presented as requirements in the source description.
## Skills & Competencies
* Distributed Systems
* Distributed Machine Learning
* Large-Scale ML Training
* Machine Learning Engineering
* Model Parallelism
* Data Parallelism
* Tensor Parallelism
* Pipeline Parallelism
* FSDP
* DeepSpeed
* Megatron
* Python
* GPU Optimization
* GPU Utilization
* Memory Management
* Large-Scale Computing
* Custom Model Sharding
* Checkpointing
* State Synchronization
* Fault Recovery
* Monitoring and Metrics
* Peer-to-Peer (P2P) Networking
* gRPC
* Routing
* NAT Traversal
* Peer Discovery
* Dynamic Routing
* Distributed Coordination
* Connection Lifecycle Management
* Network Optimization
* Latency and Bandwidth Optimization
* Production Systems
* Resilient Infrastructure
## Education & Experience
**Education:**
No specific education or degree requirement is stated in the original job description.
**Experience:**
* 5+ years of experience in distributed systems and large-scale ML training.
* Production experience building and operating distributed systems.
* Hands-on experience with distributed ML training frameworks and model-parallel training.
* Production-level Python experience.
* Experience with distributed networking and GPU workload optimization.
## Work Arrangement & Schedule
* **Location:** California, Missouri, United States
* **Work Arrangement:** Remote-first
* **Employment Type:** Senior/Staff engineering role
* **Remote Work:** Remote-first
* **Optional Office Access:** Melbourne hub
* No specific weekly hours or weekend requirements are stated in the original job description.
**Location note:** The source listing identifies the job location as **California, MO, US**, while the company description states that the role is remote-first with optional access to a **Melbourne hub**.
The source also mentions visa sponsorship for exceptional candidates and a competitive base salary for senior engineering roles in Australia.
These location-related details appear inconsistent, so the listed job location and remote-first arrangement have been preserved without inventing a correction.
## Compensation & Benefits
The original job description does not provide a specific salary range.
It states that the company offers:
* Equity-heavy compensation with meaningful ownership in a mission-driven company
* Competitive base salary for senior engineering roles in Australia
* Visa sponsorship available for exceptional candidates
* Remote-first work with optional access to the Melbourne hub
* Opportunity to work with a technical team whose members have experience at Google, Amazon, Microsoft, and leading startups
## Compliance / Additional Information
Pluralis Research describes Protocol Learning as an approach intended to support community-trained and community-owned frontier models and reduce concentration of model development, access, and economic value among a small number of large corporations.
The company is backed by **Union Square Ventures and other tier-1 investors**.