Senior Site Reliability Engineer

Laineneuralnetwork — United Arab Emirates · Posted ~2 hours ago

Senior

Skills

Site Reliability Engineering Google Cloud Platform cloud infrastructure Infrastructure as Code observability deployment pipelines production reliability incident response continuous delivery LLM APIs CI/CD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior infrastructure engineering opportunity for someone who can bridge software engineering and operations, harden production workloads on Google Cloud, build resilient pipelines for API and latency-sensitive AI workloads, and drive high availability and rapid incident resolution.

Highlights

Own the reliability roadmap for a security-sensitive AI platform, with broad responsibility across cloud infrastructure, observability, resilient production workloads, and continuous delivery.

Description

About Laine.aiLaine builds AI-powered workflows for the legal industry, where strict data security, high availability, and deterministic latency are non-negotiable. Our platform processes highly sensitive legal documents and coordinates multi-step LLM operations in real time. We are seeking a Senior Site Reliability Engineer to take ownership of our cloud infrastructure, observability, and deployment pipelines as we scale. The RoleAs our primary Site Reliability Engineer, you will bridge software engineering and infrastructure operations. You will define our reliability roadmap, harden production workloads on Google Cloud Platform, and design resilient pipelines for both traditional API services and latency-sensitive LLM workloads. You will work directly with core engineering to ensure high uptime, rapid incident resolution, and smooth continuous delivery. Key Responsibilities● Infrastructure as Code & Cloud Architecture: Design, provision, and maintain secure, reproducible cloud infrastructure across GCP (Cloud Run, Cloud SQL / AlloyDB, VPCs,...). ● Database Reliability & Scalability: Optimize PostgreSQL/AlloyDB clusters, manage connection pooling, tune query performance, implement zero-downtime database migrations with Prisma, and maintain rigorous backup/recovery strategies. ● LLM Pipeline & Service Reliability: Monitor and optimize uptime, rate limits, and latency profiles for AI model inference (Vertex AI and external LLM APIs), including tracing via LLMOps platforms (e.g., Langfuse). ● Observability & Alerting: Build centralized telemetry and tracing pipelines using OpenTelemetry, Sentry, Prometheus/Grafana, or GCP Cloud Monitoring. Establish actionable SLOs, SLIs, and alert routing. ● CI/CD & Developer Enablement: Maintain fast, reliable automated build, test, and release pipelines using Cloud Build and containerized workflows. ● Security & Compliance Posture: Enforce least-privilege IAM policies, manage secrets securely, ensure zero public exposure of sensitive databases, and maintain compliance standards required for enterprise legal technology (SOC 2, ISO 27001). ● Incident Management: Lead post-mortems, identify root causes, implement preventative safeguards, and define the team's incident escalation processes. What We Are Looking For● 4+ years of professional SRE or DevOps experience, ideally within a high-growth SaaS or startup environment. ● Deep GCP expertise: Hands-on experience configuring and debugging containerized workloads (Cloud Run), networking, IAM, and security perimeters. ● Relational Database Administration: Strong production experience tuning, monitoring, and scaling PostgreSQL or AlloyDB under heavy read/write loads. ● Software Engineering Fundamentals: Fluency with modern backend architectures, particularly TypeScript/Node.js (NestJS ecosystem) and Docker-based containerization. ● Infrastructure as Code: Proven track record managing production state and multi-environment setups. ● Production Observability: Experience architecting monitoring stacks that provide end-to-end visibility into distributed services and async job queues. ● Security-First Mindset: Practical experience locking down data layers, encryption at rest/in transit, and audit logging for sensitive enterprise customer data. On-Call & 24/7 Production SupportBecause Laine provides mission-critical infrastructure to enterprise legal teams globally, maintaining high availability and rapid incident response is critical. ● 24/7 Rotating Schedule: You will participate in a scheduled 24/7 on-call rotation shared across the engineering team, serving as a designated primary or secondary responder for production incidents. ● Response SLAs: Respond promptly to high-priority automated alerts (P1/P2) during your on-call shift, triaging critical outages, service degradations, and data pipeline failures. ● Escalation & Tooling: Leverage modern incident response tools (e.g., PagerDuty, Opsgenie, GCP Cloud Monitoring alerts) to triage issues and communicate incident status transparently. ● Sustainable Operations: We treat every page as an opportunity to fix the underlying system. You will lead blameless post-mortems, automate failure remediation, and refine alert thresholds to prevent repeat incidents and protect on-call sustainability. Nice to Have● Experience monitoring and optimizing LLM applications, token streaming latency, and vector search operations. ● Experience implementing or maintaining compliance frameworks (SOC 2 Type II, ISO27001, GDPR, FADP and equivalents). ● Prior background in legal tech, enterprise SaaS, or data-intensive contextual AI systems. Tech Stack● Cloud & Runtime: Google Cloud Platform, Cloud Run, Docker ● Data & Storage: AlloyDB / PostgreSQL, Supabase, Redis ● Backend: TypeScript, NestJS, Prisma ORM ● AI & LLM Services: Vertex AI ● Tooling & CI/CD: Cloud Build, Graphite ● Observability: Sentry, Langfuse