Description
Lead Infrastructure Engineer – Temporal Platform
Location: Remote
Job Type: Full-Time
Job Summary
We are seeking a hands-on Lead Infrastructure Engineer to build and operate the infrastructure foundation for an enterprise Temporal Platform used as an internal shared service.
The ideal candidate will have strong experience with Kubernetes/OpenShift, cloud infrastructure, SRE, Infrastructure as Code, GitOps, identity, observability, and distributed systems.
This is an opportunity to work as a key infrastructure/platform engineering team member, making critical decisions around reliability, security, scalability, and production operations.
Key Responsibilities
Build and maintain hosting, connectivity, identity, automation, observability, and production controls.Deploy and scale containerized workloads on Kubernetes/OpenShift.Configure autoscaling, networking, secrets, resource policies, and GitOps deployments.Implement enterprise SSO, IAM, role mapping, service identities, and access controls.Develop Infrastructure as Code and GitOps using Terraform, Helm, and ArgoCD.Establish monitoring and observability using Prometheus, Grafana, and OpenTelemetry.Support multi-tenant internal platforms, namespace isolation, and developer enablement.Define on-call processes and incident-response practices.Manage distributed systems and durable data, including availability, consistency, backup/recovery, retention, and performance.Partner with internal teams and vendors to establish clear operational ownership.Drive reliability, security, scalability, and production-readiness decisions.Required Qualifications
5+ years of experience in Platform Engineering, SRE, or Infrastructure Engineering.Strong hands-on experience with Cloud, Kubernetes/OpenShift, and production distributed systems.Experience with enterprise SSO/IAM and cloud identity integrations.Strong knowledge of Terraform, Helm, ArgoCD, and GitOps practices.Production experience with Prometheus, Grafana, and OpenTelemetry.Strong understanding of distributed systems, durable data, backup/recovery, and performance.Experience operating internal shared services, including multi-tenancy and namespace isolation.Hands-on experience administering or integrating Temporal, Cadence, or a comparable durable-execution/workflow platform.Proficiency in Go.Familiarity with Temporal SDKs – Java, Go, and TypeScript.Experience with enterprise networking, private connectivity, identity, and managed-service integrations.Experience establishing on-call and incident-response processes.Strong communication, troubleshooting, and independent decision-making skills.