Summary
✨ AI‑Generated
A small, senior engineering team is seeking a Staff Backend Engineer to solve difficult scalability and reliability problems in real-time AI infrastructure. You will help design and build a database from the ground up, support tens of thousands of concurrent sessions, and evolve an architecture toward substantially higher scale. The environment emphasizes technical ownership, direct collaboration, and complex distributed-systems engineering.
Highlights
Join a small, senior-heavy engineering team tackling demanding scale and reliability challenges. The role offers significant technical ownership, including designing a new database and evolving infrastructure for tens of thousands of concurrent real-time sessions in regulated enterprise environments.
Description
About the Company
Our client builds real-time voice AI infrastructure for enterprises in regulated environments, at production volume, for customers who cannot tolerate it failing.
Enterprise security and healthcare compliance are already in place.
Seed round closed, led by a well-known fund, with operator angels from several recognisable technology companies on the cap table.
The team is still in single digits.
Senior-heavy, flat, no engineering management layer, everyone reports to the founder, who holds both the CEO and CTO seat.
They were early to their market and have kept the lead through head-to-head technical evaluations.
Enterprise demand now runs ahead of what the current architecture can carry, which is why this role exists.
The Role
Our client is about to build their own database.
Not adopt one.
Build one.
They handle tens of thousands of concurrent voice sessions today and are scaling to several times that.
At that volume the off-the-shelf stack stops behaving, and the data they hold (real-time conversational data with audio, transcripts, traces and outcomes) does not sit neatly inside anything that already exists.
So they are building a purpose-built store for it.
That system is meant to be the company technical moat for the next several years.
You would lead that work, alongside owning the core backend services that keep the real-time platform standing while it scales.
This is their top-priority hire and they are making exactly one.
You report directly to the founder, so there is no layer between you and the person making the technical calls.
What is actually unsolved
What the storage layout should be when one record is an entire conversation, with audio, transcript, traces, model calls and outcomes attached, and queries run from "show me this one" to aggregating across millions.Where the boundary sits between the new store and the Postgres and ClickHouse already carrying production.What degrades first at target concurrency, and whether it degrades gracefully or simply falls over.
Nobody has proven this yet.How you make a non-deterministic pipeline debuggable, so "why did this one go wrong?" has an answer that is not an hour of log reading.
Responsibilities
Design and build the conversational data store from the ground up.
You decide what a record is, how it sits on disk, and which access patterns get to be fast.
There is no design to inherit and nobody above you to approve it.Take concurrency up close to an order of magnitude against a 99.99% target.
Nobody currently knows what breaks first.
Finding out before a customer does is the job.Own the backend services that keep real-time voice standing while all of that is happening.
TypeScript, Node.js and Python across real-time media, workflow orchestration, speech and LLM tooling.
Owning them means the pager, not just the design.Make production monitoring trustworthy at volume, so telephony events, recordings and outcomes never quietly go missing.
Silent data loss here is worse than an outage, because nobody notices it.Take "why did this conversation go wrong?" from an hour of log reading to a thirty-second answer.
OpenTelemetry and trace-first, on a pipeline where the same input can produce a different output.Ship the smallest working version in week one.
Four deploys a day is the baseline rather than a stretch goal, and designs get refined against production rather than in a document.
Requirements
The client is hiring one person into a team already in single digits and senior throughout.
The standard is not whether you clear that average.
It is whether you raise it.
These are hard filters and the process is built to test them.
You have built or core-contributed to a database, a queue or a large-scale data system.
You can name the subsystem, the decision you made inside it, and what that decision cost you.
Operating one at scale is a different thing, and the interview finds the difference quickly.You have taken something from working to working under load.
Gigabytes to terabytes a day, high cardinality, tight SLAs.
You found the bottleneck yourself, you can still quote the number before and after, and you know which of your fixes was the one that actually mattered.Internals-level with at least one of Postgres, Redis, Kafka or ClickHouse.
Internals means you know how it behaves when it is unhappy, not that you have configured it.Strong TypeScript and Node.js, plus Python.
Not languages you would be picking up here.Cloud-native infrastructure and a workflow engine.
AWS, Terraform, Kubernetes, and Temporal or Cadence in production rather than in a proof of concept.AI-native, genuinely.
Cursor, Claude Code or Codex as daily drivers, with a view on which model you reach for and where each one lets you down.
Enthusiasm with no criticism reads as shallow use.Five to fifteen years in backend or infrastructure engineering.
The process ends with a paid two to three day work trial on a real problem.
Nice to Have
Time at an observability, data infrastructure or database company.
Datadog, ClickHouse, Temporal, Stripe, Databricks, Snowflake and the tier around them.Founding engineer or early employee at a VC-backed startup, Series A or later.Real-time voice, audio pipeline or telephony systems experience.Scala, Rust or Haskell in your background.
The team reads that as systems-level thinking.Open source contributions or personal projects in systems or databases.
Culture
They deploy to production four times a day, and fixes often go out the same day a customer raises them.
Bias to action over planning.
Prototypes in days, not months.
Long whiteboard cycles do not survive.AI-native as standard.
Every frontier tool is company funded and every engineer runs them daily.
Scepticism about AI-assisted coding is a hard filter, not a preference.Full surface ownership.
Nobody has a narrow lane.
You own the parts nobody handed you.Flat by design.
No ladder, no promotion cycle, no engineering management layer.
Staff Backend Engineer
Distributed Systems and Database Infrastructure
Equity
Included
Hours
US working hours overlap required
Sponsorship
None.
US hires need citizenship or a Green Card
Experience
5 to 15 years