Site Reliability Engineer

Optomi — United States · Posted ~21 hours ago

Senior

Skills

Kubernetes Monitoring Capacity planning SLA/SLO management Infrastructure Distributed systems Infrastructure as code Automation CI/CD Troubleshooting Infrastructure as Code Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An entertainment-focused organization is seeking an experienced Site Reliability Engineer to lead reliability initiatives across critical platforms. The role emphasizes Kubernetes, scalable infrastructure, distributed systems, observability, automation, capacity planning, SLAs/SLOs, and infrastructure as code.

Highlights

Experienced SRE opportunity with technical leadership responsibilities across highly available infrastructure, Kubernetes, automation, observability, AI platforms, and distributed systems.

Description

Optomi, in partnership with a leading client in the Entertainment industry, is seeking an experienced SRE for their, Orlando, Seattle, New York, OR Burbank location. The right candidate will have Kubernetes expertise, with experience in Monitoring, capacity planning, and SLA / SLO's. The SRE will support multiple critical platforms, including a new AI platform, and an existing automation platform. Responsibilities: Support and evolve critical platforms, including an AI platform, anexisting automation platform, and a new data and observability platform.Provide technical leadership across SRE, infrastructure, platform reliability, and automation initiatives.Design, operate, and improve scalable, highly available infrastructure and distributed systems, with a strong focus on Kubernetes.Monitor system health, capacity, performance, SLAs/SLOs, and reliability; troubleshoot complex infrastructure and application issues.Build and maintain infrastructure-as-code, automation, and CI/CD capabilities using tools such as Terraform/OpenTofu, Ansible, and GitLab CI/CD.Develop reusable platform components, modules, and developer-facing tools that enable engineering teams to work more efficiently.Implement observability, security scanning, monitoring, instrumentation, and telemetry across platforms and applications.Partner across engineering, architecture, and security teams to establish technical standards and drive platform strategy.Leverage AI-assisted development and AI/ML services to improve automation, code quality, documentation, and engineering workflows. Qualifications: 5+ years of experience in SRE, software engineering, platform/infrastructure engineering, systems administration, or related fields.Deep, hands-on Kubernetes expertise is required, including operations, monitoring, capacity planning, troubleshooting, and performance management.Strong Linux administration and distributed systems experience.Experience with cloud platforms, particularly AWS, and container technologies such as Docker.Strong infrastructure-as-code experience with Terraform, OpenTofu, or similar tools.Hands-on experience with CI/CD pipelines and source control systems such as GitLab or GitHub Actions.Proficiency in at least one scripting or programming language, such as Python, Bash, Go, JavaScript, or TypeScript.Experience with monitoring, logging, and observability platforms such as Datadog, Splunk, or CloudWatch.Strong understanding of networking fundamentals, APIs, and infrastructure automation.Experience building reusable tools, platforms, libraries, or modules used by multiple engineering teams.Experience operating in large enterprise environments with cross-team collaboration and competing priorities.Strong written communication skills and the ability to produce clear technical documentation and architecture proposals.Demonstrated ability to influence technical direction and serve as a thought leader within SRE/platform engineering.Experience with AI-assisted development tools such as Cursor, Claude Code, or GitHub Copilot is preferred.Experience integrating AI/ML services into engineering or CI/CD workflows is a plus.