Software Engineer, Systems Machine Learning

Meta — United States · Posted ~22 hours ago

Mid Full-time

Skills

C++ Python Machine learning infrastructure Distributed computing ML training systems ML inference systems Performance optimization Hardware-aware optimization Profiling Instrumentation Scalable systems Low-latency systems Machine learning ML infrastructure Hardware acceleration

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Software Engineer is sought to build and optimize machine learning infrastructure operating at massive scale. The role covers distributed training and inference systems, high-performance components in C++ and Python, hardware-aware optimization, profiling, instrumentation, and low-latency execution while collaborating with research and platform engineering teams.

Highlights

Work on high-performance machine learning infrastructure at massive scale. The role spans training and inference systems, distributed computing, hardware-aware optimization, performance engineering, and collaboration with research and platform teams.

Description

Meta is seeking a Software Engineer to join our Systems ML Engineering team, focused on building and optimizing the machine learning infrastructure that powers Meta's products at massive scale. In this role, you will design and develop high-performance ML systems, working across the full stack from model training and inference pipelines to hardware-aware optimizations. You will collaborate with researchers, platform engineers, and product teams to accelerate ML workloads and improve the efficiency of AI infrastructure that serves billions of users. Software Engineer, Systems ML Responsibilities: Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelinesDevelop and maintain high-performance ML infrastructure components in C++ and Python, ensuring reliability, scalability, and low-latency executionIdentify and resolve performance bottlenecks across the ML stack using profiling, instrumentation, and benchmarking toolsArchitect and evaluate trade-offs in ML system design, including memory bandwidth, compute utilization, and I/O throughputPartner with research and product teams to translate ML model requirements into efficient infrastructure solutionsDefine and track system-level metrics and service level objectives to maintain production reliability of ML serving systemsLead technical design reviews and contribute to engineering standards for ML systems across the organizationMentor other engineers on ML infrastructure best practices, debugging methodologies, and performance optimization techniquesDrive adoption of AI-augmented development workflows to expand engineering productivity and broaden the scope of deliverablesContribute to staged rollout strategies using feature flagging and experimentation frameworks to safely deploy ML system changes Minimum Qualifications: Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience6+ years of experience in software engineering with a focus on machine learning systems, AI infrastructure, or high-performance computingExperience developing and optimizing ML training or inference pipelines using frameworks such as PyTorch, TensorFlow, or equivalentExperience with distributed computing architectures and large-scale systems design for ML workloadsExperience programming in C++ and Python for performance-critical systemsExperience using profiling and performance analysis tools to identify and resolve bottlenecks in ML or compute-intensive systems Preferred Qualifications: Experience optimizing large-scale ranking and recommendation model inference on AI accelerator hardwareExperience with hardware-software co-design, including numerics optimization and SIMD or vectorization techniquesDemonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)Experience with GPU programming using CUDA, ROCm, or equivalent hardware accelerator kernel developmentExperience with ML compiler technologies such as MLIR, LLVM, TVM, XLA, or IREEDemonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologiesExperience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) About Meta: Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics. Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment. Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com. $154,003/year to $217,000/year + bonus + equity + benefits Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.