Description
About Our Client
Our client is a rapidly growing AI infrastructure company building the systems and data infrastructure that power the next generation of AI agents.
Its platform supports frontier AI labs, Fortune 500 companies, and high-growth technology companies.
Backed by significant venture funding and experiencing strong revenue growth, our client is scaling its engineering team to meet growing demand.
The engineering team includes AI researchers, startup founders, and multiple international Olympiad medalists.
About the Role
Our client is looking for a Platform Engineer to own the reliability, scalability, performance, and developer experience of its core infrastructure and backend systems.
This is not a pure infrastructure role.
The ideal candidate combines strong production infrastructure experience with solid backend engineering skills and can reason about service architecture, APIs, databases, queues, performance, deployment safety, and production reliability.
You'll work across AWS, Kubernetes, Terraform, CI/CD, observability, and backend services to make systems more reliable, scalable, cost-efficient, and easier for engineers to build on.
Key Responsibilities
Own production uptime, latency, infrastructure provisioning, cloud costs, and incident response.Build and maintain AWS infrastructure using technologies such as Terraform, Kubernetes/EKS, Helm, Docker, EC2, ECR, S3, IAM, networking, and secrets management.Design and improve backend and platform systems for scale, including autoscaling, queueing, retries, backpressure, capacity planning, and rollback strategies.Build dashboards, alerts, logging, tracing, SLOs, runbooks, and on-call processes.Develop reliable CI/CD, release automation, environment management, and deployment workflows.Write clean, maintainable code for automation, backend services, and internal engineering tools.Identify opportunities to improve system performance, reliability, developer productivity, and cloud efficiency.
Requirements
Experience owning production cloud infrastructure for a high-availability, user-facing platform.Strong experience with AWS and containerized systems.Hands-on experience with Terraform, Kubernetes/EKS, Docker, EC2, CI/CD, networking, IAM, and cloud security.Experience operating observability, alerting, incident response, and deployment systems.Strong backend engineering fundamentals, including service architecture, APIs, databases, asynchronous systems, queues, and distributed systems.Ability to write production-quality code and apply sound engineering judgment across infrastructure and backend systems.Strong ownership mindset and ability to operate effectively in a fast-moving environment.
Nice to Have
Experience with AI/ML, data-heavy, workflow, marketplace, developer-tool, or enterprise platforms.Experience handling bursty workloads, long-running jobs, distributed workers, sandboxed execution, or high-concurrency services.Experience optimizing cloud infrastructure costs through architecture, autoscaling, workload placement, caching, cleanup, or observability.Experience building internal platforms and developer tools.Experience working in an early-stage or high-growth technology company.
Our client prioritizes technical aptitude, ownership, and learning potential over years of experience.
What Our Client Offers
Competitive compensation based on experience and location.Relocation and visa sponsorship for strong candidates moving to the US.Comprehensive medical, dental, and vision coverage.Office-provided lunch and dinner.Paid time off and company-wide holiday break.401(k) and commuter benefits.Fitness membership benefits.Generous access to leading AI development tools, including ChatGPT, Claude Code, and Cursor.Opportunity to work alongside exceptional AI researchers, engineers, and founders on cutting-edge infrastructure.