Summary
✨ AI‑Generated
A senior platform engineering opportunity to own the infrastructure and security foundations of a next-generation AI platform. The role focuses on scalable cloud systems, workload isolation, automation, and secure production environments.
Highlights
High-impact platform engineering role focused on building secure, scalable infrastructure for advanced AI workloads with strong compensation and flexible remote work.
Description
Staff Platform Engineer (GCP)
$300k Base + 20% Bonus + Stock Options
Staff Platform & Security Engineer, AI Agent Infrastructure
High-growth, venture-backed company · US Remote (hybrid optional)
About the Role
Our client is hiring a Staff-level engineer to own the infrastructure and security foundations of its AI agent platform.
As agents move from answering questions to taking real actions across production systems, the hard problems shift to identity, isolation, credentials, and auditability.
This role owns them.
This is a platform and security engineering role first.
AI is the domain; deep infrastructure and security craft is what they're hiring for.
What You'll Own
Platform & Infrastructure
Architect and operate a multi-tenant, Kubernetes-based runtime for agent workloads, including custom controllers and operators, CRDs, Helm, and per-tenant isolation.Build automation that provisions, scales, and tears down isolated environments on demand.Own infrastructure-as-code (Terraform) across cloud environments, along with CI/CD and progressive delivery.
Security & Identity
Design a security model that treats agent workloads as untrusted.Build short-lived, least-privilege credentials and delegated authorization in place of standing secrets.Own secrets management and key-managed encryption for sensitive tokens.Enforce zero-trust service-to-service communication and deny-by-default networking.Threat-model and mitigate confused-deputy attacks, privilege escalation, prompt injection, and data exfiltration.
Reliability & Observability
Run the platform as a fleet: health monitoring, automated recovery, efficient scaling, and SLOs.Build observability and audit pipelines (OpenTelemetry) that give a complete, investigable record of agent activity.Lead incident response and balance reliability, performance, and cost.
Agent Safety Controls
Build human-approval workflows for sensitive or irreversible actions.Build protected configuration that agents cannot modify.Partner with the AI team so that every agent capability is exposed through governed interfaces.
Technical Leadership
Set platform and security standards.Mentor senior engineers through architecture, not just code review.Shape the infrastructure roadmap with Product, Security, and Leadership.
What You Bring
8+ years in software, infrastructure, or distributed systems, with end-to-end ownership of security-critical production systems.Deep production Kubernetes experience: operators, controllers, multi-tenancy, workload isolation, fleet operations.Cloud-native architecture on a major cloud provider, plus strong Terraform.Applied security depth: workload identity, service mesh and mTLS, OAuth 2.0 / OIDC, token exchange, secrets management and KMS, zero-trust design.Excellent Go and/or Python.SRE fundamentals: SLOs, incident management, observability, capacity and cost engineering.High autonomy, and comfort taking ambiguous problems from architecture to production.
Nice to Have
Experience with LLM infrastructure, agent platforms, or AI security.Container sandboxing or runtime isolation technologies.
Package: competitive base, equity, bonus, full benefits.
To apply please contact Sam Shinner at Discover International