Senior DevOps Engineer

Jobright Ai — United States · Posted ~1 day ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

This role is part of the Jobright TNT - the private hiring network connecting top talent with top AI startups like Perplexity, Mercor, Cresta, Suno and 150 more. This is not a mass job posting. Only select, high-signal candidates are invited to Jobright TNT and recommended directly to hiring teams Hiring Company: Qualified Health One-liner: Qualified Health is building the healthcare-native enterprise AI platform that helps leading health systems deploy AI safely, at scale, and with measurable clinical and financial impact. Salary: $170K/yr - $220K/yr Why Join Us: • AI-native healthcare enterprise platform helping health systems safely capture the trillion-dollar healthcare AI opportunity across clinical, operational, and financial workflows • Raised $125M Series B led by NEA, now supporting 400,000+ users across leading U.S. health systems with 10x revenue growth in the past year • Founding team brings experience from Stanford Medicine, Waymo, Haven, Evolent Health, Elevance Health, UCSF, IHI, and Weill Cornell Medicine • Join a small, exceptionally high-performing team where every role directly shapes how leading health systems adopt AI safely, responsibly, and at scale Role Responsibilities • Partner with engineering teams to ensure services are production-ready before release, including reviewing deployment patterns, failure modes, resource requirements, and rollback strategies • Design and maintain observability infrastructure including metrics, logging, distributed tracing, and dashboards across multi-cloud environments • Define and manage alerting policies, SLIs/SLOs, and on-call rotations to ensure timely response to production issues • Lead and support incident response for production issues, drive root cause analysis, and coordinate hotfix deployments when needed • Author and maintain release documentation, runbooks, incident postmortems, and operational playbooks • Provide day-to-day operational support to engineering teams, unblocking deployments, debugging production issues, and improving developer experience around shipping to production • Design and maintain zero trust network architectures, ensuring secure connectivity across multi-cloud environments and tenant boundaries • Build and improve CI/CD pipelines and release processes to make production deployments safer, faster, and more predictable • Develop automation in Python and Terraform to reduce toil and codify operational best practices • Manage Kubernetes-based workloads in production, including troubleshooting cluster issues, optimizing resource utilization, and maintaining workload reliability • Operate Temporal workflows in production, including monitoring, scaling, and troubleshooting long-running workflow executions • Collaborate with security and compliance teams to maintain HIPAA and HITRUST controls across production environments Qualifications Required • 6+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, with at least 3 years directly managing production workloads • Strong proficiency with Terraform including module development, state management, and multi-environment architectures • Deep experience operating production Kubernetes environments, including troubleshooting, networking, workload management, and cluster operations • Hands-on experience with both Google Cloud Platform and Microsoft Azure services • Strong networking and security knowledge, including zero trust architectures, network segmentation, private connectivity, identity-based access controls, and secrets management • Production experience with Temporal or comparable workflow orchestration systems • Strong proficiency in Python for automation, tooling, and operational scripting • Demonstrated experience designing and operating observability stacks including metrics, logging, tracing, and alerting • Experience leading incident response, including on-call rotation management, runbook development, and postmortem processes • Track record of partnering with engineering teams to improve production readiness and release practices • Excellent written communication skills for authoring runbooks, postmortems, and release documentation • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience Preferred • Experience in healthcare industry with understanding of HIPAA compliance requirements • Familiarity with HITRUST or similar compliance frameworks • Experience operating LLM-based systems, agentic workflows, or RAG pipelines in production • Experience with GitOps workflows (Rancher Fleet, ArgoCD, or Flux) • Experience building and operating multi-tenant SaaS infrastructure • Familiarity with chaos engineering and reliability testing practices • Prior experience as a founding or early SRE/Platform hire at a startup How can I join Jobright TNT: If this is your first time applying to a Jobright TNT role, the process works as follows: 1. Apply to your first Jobright TNT role 2. We review your background to determine if you meet the TNT quality bar 3. If qualified, your application is directly recommended to the employer 4. Once accepted into TNT, you may be: - Invited to apply for other exclusive TNT-only roles - Invited to private, invite-only hiring events with top startups You will be notified of your TNT selection result. PS: All Jobright TNT roles are 100% real, directly hired by top AI startups we partner with, and come with priority review and higher response rates than the normal application queue.