Software Engineer, Distributed Systems
Features And Labels โ United States ยท Posted ~1 day ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
fal is the generative media ecosystem powering the next generation of AI products.
We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise.
For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
fal is the generative media ecosystem powering the next generation of AI products.
We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise.
For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role:
You are an experienced software engineer who thrives on building large-scale computing platforms.
You have deep expertise in large scale distributed systems that deal with high complexity, a lot of traffic and data.
You know how to achieve reliability and scale with minimum operational load.
Key Responsibilities
Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etcProduce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the worldLeverage AI to an extreme level to automate the mundane parts of building complex but reliable systemsProfile and tune low level CPU and memory performance
Requirements
3+ years experience building distributed compute and orchestration platforms in Python or RustStrong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planningDeep understanding of computational complexity and memory allocationTrack record of designing systems that scale under real production loadExperience building and using observability to drive performance and reliability decisionsExcellent communication and ability to drive technical decisions across teamsSelf-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
Experience with AI/ML inference or training infrastructureExperience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency)Background in building multi-tenant compute platformsUnderstanding of networking fundamentals and performance characteristicsFamiliarity with GPU workload characteristics and scheduling constraints
Compensation
$180,000-250,000 plus equity + benefits (This range is across all 3 levels Mid, Senior and Staff)
Location
San Francisco, CA (willing to consider remote for Senior and Staff levels)
What we offer At Fal
Interesting and challenging workA lot of learning and growth opportunitiesWe are currently hiring in downtown San Francisco.
We offer relocation assistance to San Francisco.
Health, dental, and vision insurance (US)Regular team events and offsites
We have 67,615 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 67,615 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume