Principal DevSecOps Engineer, AI Infrastructure

Comtech Global Inc — United States · Posted ~2 hours ago

Lead Full-time

Skills

Kubernetes Amazon EKS AWS Infrastructure as Code GitOps Go Python Rust Shell scripting Containerization Service networking Observability Site Reliability Engineering CI/CD Security Incident response Root cause analysis EKS Terraform Prometheus Grafana OpenTelemetry Fluentd Jaeger Backstage TypeScript

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A principal DevSecOps opportunity to build and operate the core infrastructure powering large-scale AI workloads. You will design highly available distributed systems, extend Kubernetes platforms, automate cloud environments with IaC and GitOps, establish observability and security controls, and mentor engineers in a fast-moving environment.

Highlights

Principal-level infrastructure role with ownership of highly available distributed systems and AI infrastructure. Strong emphasis on Kubernetes, cloud automation, security, observability, SRE practices, self-service platforms, and technical mentorship.

Description

Skills 8+ years operating high-availability, fault-tolerant distributed systems with IaC and GitOps.Strong coding in Go/Python/Rust plus solid shell skills; comfortable extending Kubernetes via CRDs.Deep Kubernetes/EKS expertise; mastery of containerization and service networking.Hands-on with AWS primitives (VPC, EC2, S3, IAM, RDS) and multi-region traffic/failover.Observability pro (Prometheus, Grafana, OpenTelemetry, Fluentd, Jaeger) with strong RCA/incident chops.Security fundamentals: IAM, secrets management, and compliance guardrails (SOC2/HIPAA/GDPR).Experience building secure, self-service platforms (SDKs/APIs/portals, e.g., Backstage/TypeScript).Proven SRE practice—SLIs/SLOs, error budgets—and strong testing, reviews, and CI/CD habits.Clear communicator and mentor who thrives in fast-moving environments and collaborates across ML, Responsibilities What You’ll Do Serve as the founding infrastructure engineer, building the core platform that scales the company and raises the reliability bar.Establish secure, repeatable IaC/GitOps patterns (Terraform/CloudFormation) and automated delivery (GitHub Actions, ArgoCD).Partner with teams pre-GA on design reviews, capacity planning, and readiness.Define and drive SLIs/SLOs/SLAs and an error-budget culture for services and ops.Eliminate toil with end-to-end automation across provisioning, config, testing, and operations.Co-design platforms with ML, backend, and security to safely power AI/ML workloads.Architect multi-region resilience—backup, DR, and failover—balancing availability, consistency, and cost.Advance observability and incident excellence; make smart bets on emerging infra tools.Codify production engineering standards and coach teams toward operational excellence.