Staff Site Reliability Engineer

Doghouse Recruitment — United States · Posted ~3 hours ago

Full-time Remote

Skills

Linux Kubernetes networking SRE cloud infrastructure automation observability bare metal

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A staff-level reliability engineering position focused on operating large-scale infrastructure and improving production resilience. The role involves Linux systems, container orchestration, automation, monitoring, and reducing operational complexity.

Highlights

Fully remote staff-level reliability role with significant ownership over large-scale infrastructure, automation, performance improvement, and production systems.

Description

US - 100% Remote. Staff Site Reliability Engineer – Bare Metal Linux – Data Center – Networking OTE up to $350k (base + variable), depending on experience Our client is building a cloud platform for high-throughput, compute-heavy workloads. They operate large-scale infrastructure where failure modes are real, capacity is finite, and reliability needs to be engineered, not “handled”. We’re seeking a Senior/Staff SRE who will own production reliability end-to-end for our client: define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency (p95/p99). You’ll build automation to kill toil, improve deployment safety (canary/rollback), and turn observability into signal rather than noise. This is a bare-metal environment: think Linux, datacenters, physical fleets, and real hardware constraints, not managed services. You’ll work close to the metal across Kubernetes internals (scheduling, autoscaling behavior, kubelet pressure/evictions, etcd/control plane), Linux performance (CPU/memory/I/O contention), and network debugging (DNS/TCP/TLS, packet loss, congestion). On-call is part of the job, but success is measured by how much you reduce it. Must requirements: Extensive and recent Production Engineering experience running bare metal / on-prem / data center infrastructure (not public cloud only)Deep hands-on expertise in Linux systems debugging and performance at a kernel level(CPU, memory, I/O, low-level behaviors)Strong understanding of networking (DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting under load)Strong Kubernetes experience beyond manifests: scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control planeExperience with Terraform, Docker, Helm, and modern CI/CD practicesStrong coding skills are required for this role either in Go, and/or Python, beyond automation scripting - Real engineering capability is a mustExperience in Low Latency environments. If you’re looking for complexity and a new place to nerd out on infrastructure optimization, we’d love to hear from you! Location: United States - 100% Remote. Total compensation: OTE up to $350k (base + variable), depending on experience