Site Reliability Engineer

Optomi — United States · Posted ~3 hours ago

Mid Full-time

Skills

SRE Terraform AWS observability automation performance monitoring GitLab Dynatrace

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A digital platform organization is seeking an SRE engineer to enhance application reliability through automation, cloud infrastructure, observability, and collaboration with engineering teams.

Highlights

Hands-on role improving reliability, scalability, automation, and monitoring for large-scale digital systems.

Description

Site Reliability Engineer (SRE, Terraform, AWS, Dynatrace) Optomi, in partnership with a Fortune 500 digital platform leader, is seeking a Site Reliability Engineer to join their Digital SRE team! In this role, the Site Reliability Engineer will focus on improving the reliability, scalability, and performance of critical customer-facing applications. The ideal candidate will have a strong SRE background, experience with Terraform and AWS, and a passion for observability, automation, and cross-functional collaboration. What the Right Candidate Will Enjoy! A high-impact role within a business-critical Digital SRE teamHands-on work with modern DevOps/SRE tools including Terraform, GitLab, and DynatraceCollaborative environment interacting with application teams, sustain partners, and leadershipExposure to large-scale systems architecture and service level objective (SLO) managementOpportunity to contribute to automation initiatives and performance monitoring best practicesExperience of the Right Candidate: 3+ years of experience in a Site Reliability Engineer, DevOps, or similar infrastructure-focused roleStrong experience with Terraform and GitLabHands-on experience with observability and monitoring tools, preferably Dynatrace and SplunkBackground in AWS or similar public cloud platformsExperience working with or supporting APIs and/or customer-facing front-end applicationsProven ability to lead incident response efforts and create effective runbooks and postmortemsFamiliarity with service level indicators (SLIs), error budgets, and implementing SLO frameworksEffective communicator and collaborator who thrives in a cross-functional environmentResponsibilities of the Right Candidate: Monitor and manage production environments with a focus on reliability and performanceProvide on-call support on a rotating basis, ensuring rapid response to production issuesImplement and maintain observability standards across systems using tools like DynatraceBuild automation to eliminate manual operational work and streamline triage processesLead service readiness efforts including documentation, health scoring, and SLO adoptionCollaborate with sustain partners and application teams to resolve issues and drive stabilityParticipate in blameless postmortems and promote continuous improvement practicesSupport cross-functional initiatives to improve platform performance and uptime