Description
About The Company
Hims & Hers is a leading health and wellness platform dedicated to transforming the way healthcare is delivered.
Our mission is to help individuals feel their best by providing accessible, affordable, and personalized health solutions.
We believe that everyone deserves to have access to quality care that is tailored to their unique needs, from diagnosis through treatment and ongoing management.
By normalizing health and wellness challenges and innovating on their solutions, we aim to make better health outcomes easier to achieve for all.
As a publicly traded company on the NYSE under the ticker symbol “HIMS,” we are committed to creating a culture of innovation, inclusivity, and excellence.
To learn more about our offerings and company culture, visit hims.com/about and hims.com/how-it-works.
Our focus on talent-first, flexible, and remote work environments underscores our dedication to employee well-being and professional growth.
About The Role
We are seeking a highly skilled Staff ML Systems Engineer to join our team and take ownership of designing, building, and maintaining the production infrastructure that powers our AI initiatives.
This role is deeply technical and hands-on, focusing on the systems that support AI and machine learning operations within our organization.
You will be responsible for managing critical infrastructure components such as Kubernetes platforms, CI/CD pipelines, infrastructure-as-code, inference and model-serving infrastructure, and observability and tracing systems.
Your work will ensure that AI services are reliable, secure, and compliant within our regulated healthcare environment.
You will collaborate closely with ML engineers, product teams, clinical stakeholders, and security teams to create a robust, scalable, and efficient AI infrastructure.
This position offers an exciting opportunity to influence how AI technology is operationalized at a company where it directly impacts patient outcomes and healthcare delivery.
Qualifications
8+ years of professional experience in infrastructure, platform, DevOps, or SRE engineering, with at least 3 years focused on ML/AI systems in production.Deep, hands-on experience with Kubernetes, preferably EKS, including autoscaling, cluster management, and process orchestration.Strong skills in infrastructure-as-code tools such as Terraform, and designing secure cloud architectures including IAM, OIDC, and secrets management.Proficiency in Python for building production infrastructure tooling, CLIs, and observability pipelines.At least 2 years of experience operating LLM-based systems in production, including inference routing, serving, and reliability patterns.Experience with observability and tracing stacks such as Datadog, OpenTelemetry, Langfuse, and related metrics and log pipelines.Expertise in designing and maintaining CI/CD pipelines, build systems, and developer tooling for rapid deployment cycles.Strong systems-thinking mindset with a focus on failure modes, SLOs, security, and long-term maintainability.Proven ability to write and lead technical design documents, RFCs, and infrastructure standards.Excellent collaboration skills across engineering, ML, product, security, and clinical teams.Experience working in regulated environments such as healthcare, fintech, or life sciences is preferred.
Responsibilities
Own and scale the AI compute and deployment platform, including cluster operations, autoscaling, storage, and workload isolation.Develop and maintain GitOps-based deployment pipelines using Helm/Kustomize to enable safe, repeatable AI service deployment.Design episodic and preview environments, feature-branch deployments, and nightly release pipelines for validation and testing of AI models.Build and operate inference infrastructure and multi-provider LLM AI gateways, managing credentials, rate limits, and failover mechanisms.Create reliable serving patterns for LLM workflows, including routing, grounding, tool execution, and context assembly.Develop reusable infrastructure abstractions and contracts to standardize AI service deployment and consumption across teams.Manage the observability and tracing stack, ensuring AI behavior is auditable, debuggable, and reliable in production.Build analytics and monitoring pipelines to surface latency, error rates, quality metrics, and regression signals to relevant stakeholders.Define SLOs, incident response procedures, and on-call runbooks for AI infrastructure, leading troubleshooting efforts to improve reliability.Enhance the AI developer platform and CI/CD pipelines, reducing build and deployment cycle times to increase developer velocity.Implement security, compliance, and governance measures, including IAM, secrets management, and privacy controls, to ensure HIPAA compliance.Drive strategic infrastructure initiatives, including cluster architecture, inference platform development, GPU compute strategies, and observability improvements.Write technical design documents, lead reviews, and set standards for infrastructure and development workflows.Mentor engineering teams on reliability engineering, infrastructure-as-code, and MLOps best practices, facilitating a smooth transition from prototypes to production systems.
Benefits
Competitive salary and equity compensation packages.Unlimited PTO, company holidays, and quarterly mental health days to promote work-life balance.Comprehensive health benefits including medical, dental, vision, and parental leave.Employee Stock Purchase Program (ESPP) and 401(k) plan with employer matching contributions.Opportunities for offsite team retreats and professional development.A flexible, remote work environment that supports your personal and professional growth.
Equal Opportunity
We are proud to be an Equal Employment Opportunity employer.
We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, or any other basis protected by federal, state, or local law.