Site Reliability Engineer
Bam Ventures Llc — United States · Posted ~11 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
What is Aisle?
Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened.
Aisle builds a real-time control layer on top of physical retail.
We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact.
This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time.
Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar.
Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment.
We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis.
You'll win here if…
You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actioCan solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the systemYou’re comfortable moving quickly, shipping improvements, and iterating in productionAI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems fasterYou treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc.
(gist of harness engineering)You have experience building or orchestrating AI/agent workflows - in production or through serious side projects
About Your Role
Reliability & Infrastructure
Own the reliability, scalability, and observability of our infrastructure across GCP and VercelDesign and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboardsManage IAM, service accounts, and security best practices across our cloud environmentParticipate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements
Core Infrastructure & State
Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflowsBuild and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deploymentsInvestigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolutionStabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)Automate infrastructure provisioning, deployments, and operational workflows
AI & Next-Gen Tooling: Agent Ops
Build agent operations infrastructure that enables AI agents to run safely and reliably in productionDevelop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetryHelp define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loopOwn visibility into AI usage, reliability, and spend as our agent footprint scales
Cross-functional Impact
Partner closely with engineering and product teams to maintain reliability without slowing development velocityAct as a force multiplier across the team — helping engineers ship faster and more safely
About your skillsMust haves
4+ years in SRE, DevOps, or infrastructure/platform engineeringStrong, hands-on experience with a major cloud platform (preferable GCP)Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)Solid understanding of IAM, security, and cloud best practicesExperience with observability tools like DatadogFamiliarity with Node.js environmentsAI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.Experience building or orchestrating AI/agent workflows (work or serious personal projects)High ownership, strong curiosity, and a bias toward action
Nice to haves
Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controlsFamiliarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plusHands-on experience with GCP Workflows for orchestrationFamiliarity with Prisma, pgbouncer, and PostgreSQL connection poolingExperience with Vercel deployment and edge computingFamiliarity with the k8s ecosystemFamiliarity with Redis and BullMQUnderstanding of SOC 2 compliance requirements and implementationPrevious experience in a high-growth startup environmentPrevious backend engineering experience to bridge the gap between infrastructure and code
About the stack
Cloud: Google Cloud Platform (GCP)Infrastructure: Cloud Functions, Pub/Sub, Workflows, KubernetesDatabase: PostgreSQL with pgbouncer, PrismaObservability: DatadogRuntime: Node.js, TypeScriptDeployment: Vercel, GCPFrontend: React, Next.js, TypeScriptBackend: TypeScript, PostgreSQL, Next.js
We have 113,417 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 113,417 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume