Skills
Kubernetes
Amazon EKS
AWS
Infrastructure as Code
GitOps
Go
Python
Rust
Shell scripting
Containerization
Service networking
Observability
Site Reliability Engineering
CI/CD
Security
Incident response
Root cause analysis
EKS
Terraform
Prometheus
Grafana
OpenTelemetry
Fluentd
Jaeger
Backstage
TypeScript
Summary
✨ AI‑Generated
A principal DevSecOps opportunity to build and operate the core infrastructure powering large-scale AI workloads. You will design highly available distributed systems, extend Kubernetes platforms, automate cloud environments with IaC and GitOps, establish observability and security controls, and mentor engineers in a fast-moving environment.
Highlights
Principal-level infrastructure role with ownership of highly available distributed systems and AI infrastructure. Strong emphasis on Kubernetes, cloud automation, security, observability, SRE practices, self-service platforms, and technical mentorship.
Description
Skills
8+ years operating high-availability, fault-tolerant distributed systems with IaC and GitOps.Strong coding in Go/Python/Rust plus solid shell skills; comfortable extending Kubernetes via CRDs.Deep Kubernetes/EKS expertise; mastery of containerization and service networking.Hands-on with AWS primitives (VPC, EC2, S3, IAM, RDS) and multi-region traffic/failover.Observability pro (Prometheus, Grafana, OpenTelemetry, Fluentd, Jaeger) with strong RCA/incident chops.Security fundamentals: IAM, secrets management, and compliance guardrails (SOC2/HIPAA/GDPR).Experience building secure, self-service platforms (SDKs/APIs/portals, e.g., Backstage/TypeScript).Proven SRE practice—SLIs/SLOs, error budgets—and strong testing, reviews, and CI/CD habits.Clear communicator and mentor who thrives in fast-moving environments and collaborates across ML,
Responsibilities
What You’ll Do
Serve as the founding infrastructure engineer, building the core platform that scales the company and raises the reliability bar.Establish secure, repeatable IaC/GitOps patterns (Terraform/CloudFormation) and automated delivery (GitHub Actions, ArgoCD).Partner with teams pre-GA on design reviews, capacity planning, and readiness.Define and drive SLIs/SLOs/SLAs and an error-budget culture for services and ops.Eliminate toil with end-to-end automation across provisioning, config, testing, and operations.Co-design platforms with ML, backend, and security to safely power AI/ML workloads.Architect multi-region resilience—backup, DR, and failover—balancing availability, consistency, and cost.Advance observability and incident excellence; make smart bets on emerging infra tools.Codify production engineering standards and coach teams toward operational excellence.