Site Reliability Engineer

Bam Ventures Llc — United States · Posted ~11 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

What is Aisle? Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened. Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact. This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time. Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment. We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis. You'll win here if… You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actioCan solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the systemYou’re comfortable moving quickly, shipping improvements, and iterating in productionAI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems fasterYou treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering)You have experience building or orchestrating AI/agent workflows - in production or through serious side projects About Your Role Reliability & Infrastructure Own the reliability, scalability, and observability of our infrastructure across GCP and VercelDesign and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboardsManage IAM, service accounts, and security best practices across our cloud environmentParticipate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements Core Infrastructure & State Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflowsBuild and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deploymentsInvestigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolutionStabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)Automate infrastructure provisioning, deployments, and operational workflows AI & Next-Gen Tooling: Agent Ops Build agent operations infrastructure that enables AI agents to run safely and reliably in productionDevelop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetryHelp define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loopOwn visibility into AI usage, reliability, and spend as our agent footprint scales Cross-functional Impact Partner closely with engineering and product teams to maintain reliability without slowing development velocityAct as a force multiplier across the team — helping engineers ship faster and more safely About your skillsMust haves 4+ years in SRE, DevOps, or infrastructure/platform engineeringStrong, hands-on experience with a major cloud platform (preferable GCP)Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)Solid understanding of IAM, security, and cloud best practicesExperience with observability tools like DatadogFamiliarity with Node.js environmentsAI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.Experience building or orchestrating AI/agent workflows (work or serious personal projects)High ownership, strong curiosity, and a bias toward action Nice to haves Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controlsFamiliarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plusHands-on experience with GCP Workflows for orchestrationFamiliarity with Prisma, pgbouncer, and PostgreSQL connection poolingExperience with Vercel deployment and edge computingFamiliarity with the k8s ecosystemFamiliarity with Redis and BullMQUnderstanding of SOC 2 compliance requirements and implementationPrevious experience in a high-growth startup environmentPrevious backend engineering experience to bridge the gap between infrastructure and code About the stack Cloud: Google Cloud Platform (GCP)Infrastructure: Cloud Functions, Pub/Sub, Workflows, KubernetesDatabase: PostgreSQL with pgbouncer, PrismaObservability: DatadogRuntime: Node.js, TypeScriptDeployment: Vercel, GCPFrontend: React, Next.js, TypeScriptBackend: TypeScript, PostgreSQL, Next.js