Description
Staff Cloud Engineer
Remote (U.S.
& Canada)
$210K–$260K base + equity
Sage Recruiting is partnering with an AI-native, early-stage infrastructure startup that's tackling a problem every engineering leader is quietly panicking about: AI coding agents now write code faster than any human team can review it, and the old ways of enforcing standards (wikis, checklists, manual review) were never built for that volume.
Our client has built a guardrails engine that turns a company's engineering standards into automated enforcement, on every commit, every pull request, and every deploy, for human and AI-written code alike.
They're small, senior, and heavily AI-leveraged, moving with the speed and focus of a team several times their size.
They've already landed their first paying customers and have real momentum building.
Top-tier venture investors and an angel bench of well-known operators and creators from the developer tools world back them.
This is a company at an inflection point between "early traction" and "real scale," actively going live with large enterprise customers, and this hire is one of the people who will build the infrastructure that the next chapter runs on.
The Role
This is a staff-level, deeply technical individual contributor role for a true backend/cloud/ops generalist.
You'll be handed meaty, ambiguous problems and trusted to own them end-to-end: design, build, ship, and fix, with a seat in the on-call rotation like everyone else on the team.
You'll write and own production Go code for the core platform, and you'll design and run the AWS and Kubernetes infrastructure it lives on, across both a hosted service and customer-managed deployments.
A big part of the job is taking early, minimal product surfaces and making them enterprise-ready: authentication, backup, availability, and the reliability bar that large customers expect.
You'll stay hands-on while helping decide what gets built and how, and you'll work directly with the company's forward-deployed engineers and customer platform teams on hard deployment and scale problems, feeding what you learn back into the product.
Who You Are:
You're a senior Platform Engineer who's comfortable working in code and cloud infrastructure.You're comfortable being handed an ambiguous, high-stakes problem and can be trusted to run with it end-to-end without a fleshed-out playbook.You think like a startup engineer: you know where the smart trade-offs are, where it's fine to cut a corner, and where it isn't, rather than defaulting to the most complete or "correct" solution.You lean Kubernetes-strong in particular.You're AI-native in your daily workflow, and you're excited by the idea of taking a product that's complete but still early and making it ready for large enterprise customers.
What You'll Do
Write and maintain production Go code in the backend and supporting services, owning features from design and testing through deployment and ongoing operationDesign and evolve the AWS architecture and Kubernetes infrastructure, making deliberate trade-offs around reliability, security, performance, cost, and how much complexity a small team can supportBuild reusable Terraform modules, Helm charts, and deployment tooling that make provisioning, upgrades, and recovery repeatable, for the hosted service and customer-managed installations alikeHarden early-stage product surfaces to meet enterprise expectations: authentication, backups, availability, and the operational maturity large customers requireImprove production reliability: useful metrics and alerts, capacity planning, backups and tested recovery, safe rollouts, and clear rollback paths.
Take your turn in the on-call rotation and fix the root causes of recurring problems, not just the symptomsDebug across the full stack, following a failure through Go code, a database query, container behaviour, Kubernetes networking, or cloud infrastructure rather than stopping at a team boundaryImprove CI/CD and the development environment so engineers can test realistic changes and ship frequently without making production fragilePartner with forward deployed engineers and customer platform teams on difficult deployment and scale problems, feeding what's learned back into the productLead technical design and code reviews, mentor teammates, and document decisions well enough that other engineers can maintain what you buildBuild and manage AI-assisted engineering workflows for implementation, review, testing, and operational investigation, with appropriate access controls and checks on their output
You must have:
Strong Go and software engineering fundamentals: production services, concurrency, APIs, testing, performance work, and debugging distributed systems.
This doesn't need to be your single deepest specialty, but you're comfortable owning backend code end-to-endDeep cloud experience, ideally AWS (we're open to strong GCP or Azure backgrounds too): networking, IAM, compute, storage, and managed databases, with a real understanding of failure modes, isolation boundaries, and costExtensive, strong Kubernetes experience: built and operated production clusters and workloads, handled upgrades, debugged real failures, and worked with scheduling, networking, storage, RBAC, resource management, and Helm beyond an install guide.
This is where we most want depthExtensive Terraform experience: reusable modules, state, environment separation, drift, and safe changes to existing production infrastructureStrong operational instincts: comfortable with Linux, Docker, networking, and troubleshooting under pressure, and you've owned systems after launch and made them easier to operate over timeWorking knowledge of production data stores and observability: you can operate a SQL database (PostgreSQL or comparable) and object storage like S3, and you know how to instrument and read metrics, logs, and traces (Prometheus, Grafana, OpenTelemetry) well enough to actually run a system, not just build oneExtensive, hands-on use of AI coding tools and agents as part of your daily work.
You build your own workflows, supply the context and tools they need, manage their permissions/cost/failure modes, and can explain how you verify what they produceA track record of owning ambiguous, critical problems end-to-end and delivering changes that held up in production, whether that's 7+ years as a staff-level IC or a faster growth trajectory that's gotten you thereClear communication and high ownership: you write useful design notes, can explain a trade-off to another engineer or a customer, and make progress without waiting for a detailed spec
NOTE: This isn't a fit for someone who wants an architecture or management role with little hands-on coding, or who enjoys operating infrastructure but doesn't want to build production software.
It's also not a fit for someone who uses AI only occasionally or needs a narrowly defined remit, with a separate team handling everything outside it.
It's a bonus if you have:
Something that makes you stand out: a well-known open-source project you've built or maintained with real traction, ownership of a critical production system early in your career, multi-disciplinary range (ops plus frontend, for example), or a background at a company known for excellent engineering craftExperience building developer tools, infrastructure products, or other B2B software used by engineering teamsExperience shipping software into customer-managed Kubernetes environments, including restricted networks and enterprise security requirementsgRPC and Protocol Buffers, queues/background workers, or distributed job executionGitOps and progressive delivery, Kubernetes operators, or multi-region systemsFamiliarity with software supply-chain security, policy as code, or compliance automationA founder or early-stage startup background, or a mix of big-tech and startup experience
What Success Looks Like
First 30 days: you're getting familiar with as much of the system as possible, shipping small, complete contributions across most areas of the product to build context fast.By 60 days: you're taking on larger, more substantial projects.By 90 days: you're consistently taking on real work and shipping it end-to-end without oversight.Six months in: you're a trusted, independent owner of critical parts of the system, someone the rest of the team can hand hard problems to.
Why Join
Get in early with a company at a genuine inflection point, actively going live with large enterprise customers and scaling fastEvery team member has a real voice in product direction and execution, not just their own corner of the codebaseYou'll be supporting the developer infrastructure of large enterprises: your work directly shapes the experience of other engineers, including platform teams at big, complex organizationsReal architectural ownership from day one, working directly with the founder and founding engineers, with room to shape both the product and how the team builds itA small, senior team moving fast, with heavy investment in AI tooling and none of the bureaucracy of a bigger companyGrowth here is squarely on the IC track: real room to grow your technical scope and seniorityFully remote, with the team kept within a few hours of each other, so there are no late-night meetings
Compensation & Benefits
$210,000–$260,000 base plus equity
401(k)/RRSP matching
Healthcare, dental, and vision, including dependents
Life insurance
Health/Lifestyle stipend
Fully remote work