AI/ML Infrastructure Engineer

Predii β€” United States Β· Posted ~4 hours ago

Mid Full-time Remote

Skills

Machine Learning infrastructure Cloud infrastructure Python Data platforms Distributed systems MLOps Cloud Machine Learning Distributed Systems Docker Kubernetes

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A growing AI-focused engineering team is looking for an infrastructure specialist to build reliable platforms supporting machine learning workloads, large-scale data processing, and intelligent applications.

Highlights

Work on large-scale AI infrastructure, advanced data systems, and high-impact engineering projects with opportunities for research and innovation.

Description

AI/ML Infrastructure Engineer Mid-Level to Senior | Engineering & Platform Ops US-based β€” CA preferred, open to West Coast + remote. Hybrid-friendly. Address: 2211 Park Blvd, Palo Alto, CA 94306 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ ABOUT PREDII━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence β€” powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency. We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships β€” and real customers use it. Learn more at www.predii.com. PREDII RESEARCH━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting "confident-but-wrong" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content. We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead. And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin. THE VIBE━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ We need a AI/ML Infrastructure Engineer who wants more than tickets β€” someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You'll touch: Multi-cloud infra (Azure, GCP, AWS)Kubernetes, CI/CD, automation-everythingSecurity, compliance, access β€” keeping the house lockedMonitoring & reliability β€” catching problems before customers doIncident response & disaster recoveryCloud cost optimization (yes, we care about the bill too)Senior folks: expect to shape architecture and mentor the team, not just execute someone else's roadmap. WHAT YOU'LL ACTUALLY DO━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Cloud & Platform β€” build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.DevOps & CI/CD β€” ship pipelines that are reliable and repeatable, not held together with duct tape.Reliability & Observability β€” build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.Security & Compliance β€” RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.Resilience & Ops β€” own backups, DR, capacity planning, cloud costs, and production support.Keep Leveling Up β€” evaluate new tools across DevOps, infra, and DevSecOps; you're not stuck with 2019's stack.IT Support β€” jump in on Windows Server / macOS support when needed. YOUR TOOLKIT━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Cloud & Containers: Azure, GCP, AWS, Docker, Kubernetes, AKSAutomation: Terraform, Ansible, Helm, Bash, PythonCI/CD: GitHub Actions, GitLab CI, Azure DevOps (or similar)Observability: Grafana, Prometheus, ELK (or equivalent)Networking: TCP/IP, DNS, load balancing, VPNs, firewalls/NSGs, segmentationIdentity: RBAC, cloud IAM, Okta, multi-tenant authSystems: Linux, Windows Server, macOSSecurity & Compliance: SOC 2 Type 2, access/change controls, business continuity + DR WHAT YOU BRING━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 3–5+ years hands-on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production β€” not just labs.Strong cloud + Kubernetes chops β€” Azure/AKS preferred; GCP/AWS/Docker is a big plus.Solid networking and security fundamentals.Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).Real Git-based CI/CD experience β€” automated deploys, security scanning included.Battle-tested on observability & prod ops β€” monitoring, incident response, RCA, runbooks, backup, DR.Working knowledge of security/compliance across Linux, Windows Server, macOS.A self-starter mindset β€” comfortable working independently across a distributed US–India team. BONUS POINTS━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Auth/IAM experience, especially multi-tenant setups.Background in automotive or data-heavy platforms.Been through a SOC 2 Type 2 audit before.Startup or small-team energy β€” you've worn more than one hat.Cloud, Kubernetes, or Terraform certs.Senior folks: mentoring or technical leadership experience. HOW WE ROLL━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Own it β€” flag issues early, make the call, follow through.Be proactive β€” don't wait to be told; spot the risk, bring the fix.Share the load β€” security and reliability are everyone's job, not just yours.Make it count β€” your work ships to production and touches real customers, real fast. How To ApplyEmail: jobs@predii.com