Site Reliability Engineer

Workistry — United States · Posted ~1 day ago

Senior Full-time

Skills

Site Reliability Engineering Software Engineering Production Operations Database Optimization Application Performance Optimization Kubernetes Helm CI/CD TypeScript Python ML Deployment ML Databases

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology company is seeking a Site Reliability Engineer to own production infrastructure and improve application reliability and developer experience. The role combines software and operations engineering, with responsibility for large-scale machine deployments, Kubernetes and Helm, CI/CD optimization, database access, production troubleshooting, and performance improvements. Strong Python and TypeScript experience is valuable for supporting high-velocity, reliable deployments.

Highlights

Own a large-scale production environment, work across infrastructure and application reliability, optimize deployment pipelines, and improve engineering productivity and system performance.

Description

Kaystaff is hiring a Site Reliability Engineer for a Top Technology Company In New York and SF. This position requires a balance of operational and software engineering skills. This role involves optimizing database accesses, handling production issues, and improving application performance. As our SRE, you will be responsible for our entire production environment and improve the development experience across both infrastructure and application reliability. Infrastructure Responsibilities Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high‑velocity feature release while maintaining the highest reliability.DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform. Requirements 7+ years of experience as a highly technical, application‑ leaning Site Reliability Engineer with prior high‑ growth start‑ up and software engineering experience.Experience driving systematic application improvements across complex, distributed systems with high‑availability requirements for a high‑growth start‑upExtended period working at high‑ growth start‑ ups (no recent Big Tech)Prior SWE experience and strong coding abilities (Python and TypeScript)Ability to own complex systems and solve complex problems independentlyDefines standards for operational excellence, e. g. , automating manual workflows Salary ‑ $200k ‑ $300k/yr.