Summary
✨ AI‑Generated
A senior Platform DevOps Engineer is sought to deploy and operate a Kubernetes-based healthcare analytics platform in a new cloud or on-premises environment. The role involves integrating Kubernetes clusters, configuring networking and security, managing certificates and DNS, using infrastructure automation, and resolving complex deployment issues. Strong hands-on Kubernetes and infrastructure engineering expertise is essential.
Highlights
Hands-on senior engineering role focused on deploying and operating a sophisticated healthcare AI platform, with substantial ownership of Kubernetes infrastructure, automation, networking, security, and troubleshooting across cloud or on-premises environments.
Description
Senior Platform/DevOps Engineer (Kubernetes)
Healthcare AI Platform Deployment
Role Mission
Deploy and operate a prebuilt, Kubernetes-based healthcare analytics/AI platform in a new cloud or on-prem environment using the existing automation (Ansible, Make, Helm).
This role requires a deep technical understanding of the codebase's architecture to ensure correct configuration and troubleshoot complex deployment issues.
Core Responsibilities
The successful candidate will be a hands-on executor focused on standing up one new environment (cloud or on-prem) with repeatable redeployments.
· Kubernetes Cluster Integration: Bring up, configure, and integrate a new managed (e.g., EKS) or self-hosted Kubernetes cluster to serve as the foundation for the platform.
· Infrastructure Configuration: Adapt and configure all necessary networking and security components:
o Ingress controllers and certificates (based on certificates.yml, ingress-controller, and DNS requirements).
o DNS and L4/L7 load balancers for internal and external access.
· System Component Integration: Configure and integrate essential platform services, including:
o Secrets management with Vault.
o Identity Provider (IdP) integration, likely Keycloak/OIDC, for user authentication and access controls.
o Object storage configuration (MinIO/S3) for the Data Lake.
· Build & Deployment Pipeline:
o Execute existing build processes (Makefile, containers.yml, charts.yml) to build and push platform component images and Helm charts to container registries.
o Set up and maintain the CI/CD pipelines (e.g., *.gitlab-ci.yml) for automated deployment.
· Codebase and Logic Interpretation:
o Read and interpret the platform’s Python/FastAPI services, Helm charts, and Ansible playbooks to validate system behavior and troubleshoot deployment failures.
o Deeply understand what each platform component (api, repository, vdi, data-lake, compute-cluster) does and how they are orchestrated.
· System Hardening: Tune storage classes, networking, and security policies, and perform necessary smoke tests and hardening steps before handoff.
· Documentation & Handoff: Document all environment-specific steps, configuration details, and create clear operational runbooks and handoff guides.
Must-Have Technical Skills
· Kubernetes Operations (Expert): Strong operational experience with K8s cluster lifecycle, deploying complex multi-component applications, and using kubectl.
· Infrastructure-as-Code (IaC): Proficiency with Terraform for provisioning cloud/on-prem infrastructure and Ansible for configuration management and playbook execution.
· Container and Artifact Management: Expertise with container registries, image building, Docker, and the image lifecycle.
· Helm: Strong experience with Helm for packaging and deploying Kubernetes applications, including writing and debugging custom charts.
· Networking: In-depth knowledge of DNS/TLS, ingress controllers, L4/L7 load balancing, and firewall rules in a cloud and Kubernetes context.
· Linux & Shell Scripting: Proficient in Linux system administration and Bash/Shell scripting for automation and diagnostics.
· Security Components: Hands-on experience with Vault and setting up Keycloak/OIDC or other identity providers.
· Storage: Experience with object storage (S3/MinIO) and configuring Kubernetes persistent volumes (PVs) and storage classes.
· Code Literacy: Proven ability to read and reason about Python/FastAPI application code, as well as complex YAML/Helm/Ansible configurations, to validate and troubleshoot system behavior.
Nice-to-Have Skills
· VDI/remote desktop solutions on Kubernetes (e.g., using VNC or similar technologies as seen in the codebase).
· Experience setting up observability (monitoring, logging, tracing).
· Experience in regulated environments, such as healthcare or finance.
· Ability to implement small fixes or diagnostics in Python/FastAPI.
Working Style
· A hands-on, bias-for-action executor who prioritizes reusing and adapting existing playbooks and Makefiles over rewriting.
· Proactive and meticulous in identifying and addressing environment-specific constraints and dependencies.
· Clear and concise documentation skills.
· Collaborates effectively with AI/ML tools like Cursor and OpenAI Codex for rapid troubleshooting and code adaptation.
Language
Fluent in English (written and verbal) for clear documentation and cross-functional collaboration.