Description
About the Company
A venture-backed software company rebuilding the systems US mortgage lenders run on.
The category is loan origination: the software a lender uses to move a mortgage from application to funded.
The products in market were built twenty years ago, they are slow and unstable, and the human cost of originating a single loan is still over $11,000.
Tens of billions of dollars in new home loans have run through their platform since 2020.
Under fifty people, remote-first across the United States since the company started, engineering-led.
Recently funded, with three years of runway.
A large US mortgage lender goes live on the platform this year, the engineering roadmap is organised around that date, and on the company’s own account the go-live makes the business profitable.
The stated ambition beyond that date is fully automated origination.
On the company’s own numbers, that would take over $16 billion a year out of lender operating expense.
The Role
You would build the platform under the product, not the product itself.
The AWS estate and everything that runs on it: the infrastructure in Terraform and Pulumi, the internal developer tooling and shared TypeScript libraries, the CI/CD pipelines and the ephemeral environment that stands up for every pull request, and the observability underneath all of it.
Infrastructure treated as a product, whose customers are the engineers shipping the product itself.
The stack: AWS underneath everything: ECS, Lambda, VPC, ALB, IAM, RDS with Aurora PostgreSQL, ElastiCache, MSK, OpenSearch, S3, CloudWatch, CloudTrail, GuardDuty.
Infrastructure as code in Terraform and Pulumi.
CI/CD in Buildkite.
TypeScript throughout the product: React at the front, Node and Express in Docker on the backend, GraphQL and Apollo on both sides.
Observability in Datadog and Sentry, Cloudflare at the edge.
What they have already built for engineers: TypeScript types generated from the Postgres and GraphQL schemas, so a database change surfaces as a type error in the React components it breaks.
Code shared between client, server and other packages in one monorepo.
A temporary staging environment deployed for every pull request.
A library of test utilities that mocks production-shaped data.
This role owns that platform and decides what it becomes.
Responsibilities
Own and evolve the AWS infrastructure in Terraform and Pulumi, run as a product whose customers are the engineering team.Design and maintain the internal developer tooling: shared TypeScript libraries, SDKs and code generation that standardise how software is built and shipped.Build and maintain the CI/CD pipelines in Buildkite and the per-pull-request ephemeral environments that make deploying feel easy and safe.Drive reliability through SLOs, autoscaling, incident response and postmortems, with observability in Datadog.Own the reliability and performance of Aurora PostgreSQL as enterprise load arrives.Enforce security across IAM, secrets management, encryption, audit logging and DDoS protection, in a product that holds some of the most sensitive personal data people have.Reduce toil through automation and self-service tooling, inside a shared on-call rotation.
Requirements
You have built and run platform, infrastructure or developer tooling at a high-growth startup where software was the product.
A career spent only inside Big Tech does not fit; Big Tech followed by a startup does.You have built and maintained CI/CD pipelines and per-pull-request ephemeral environments, and you can say what they made safe and what they cost to keep fast.Production TypeScript and Node.
You may not ship product features every day, but you read, navigate and change a large TypeScript codebase comfortably, and you have built backend libraries or internal tooling that other engineers depend on.Deep AWS across compute, networking, storage and security: ECS, Lambda, VPC, ALB, IAM, RDS, ElastiCache, MSK, OpenSearch, S3, CloudWatch, CloudTrail, GuardDuty.Strong Terraform and/or Pulumi: modules, workspaces, and CI-driven plan and apply workflows.Reliability engineering you have run in production: SLOs, error budgets, incident response.Seven to twenty-five years as an engineer, with the recent years focused on platform work: internal developer tooling and AWS infrastructure.You write clearly enough that other people can operate what you built.
Nice to Have
A track record building production observability stacks: Datadog, CloudWatch, Sentry, distributed tracing, SLOs.Internal platforms that measurably sped engineers up: code generation, CLIs, templates, shared SDKs, frameworks.Curiosity about what AI and LLM workloads demand from infrastructure.An active GitHub, or public work you can point at.