Senior Site Reliability Engineer

Nextgenit Inc — United States · Posted ~21 hours ago

Senior Hybrid

Skills

Site Reliability Engineering DevOps Incident management Blameless post-mortems SLOs SLIs Error budgets Monitoring Alerting Incident response Problem management Automation Software engineering IT operations SRE SLO SLI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior Site Reliability Engineer is sought to strengthen the reliability, availability, performance, and operability of software systems. The role combines software engineering with IT operations, emphasizing automation, monitoring, alerting, incident response, problem management, and reduction of operational toil. You will help define reliability standards and production-readiness requirements while partnering with engineering, architecture, and product teams in a fast-paced DevOps environment. The position follows a hybrid work model.

Highlights

Hybrid senior SRE role focused on improving system reliability, availability, performance, and operational efficiency. Offers the opportunity to automate operational work, establish strong monitoring and incident-response practices, and collaborate closely with engineering, architecture, and product teams.

Description

Job Title: Sr. Site Reliability Engineer (SRE) Location: Kalamazoo, MI (3 days a week hybrid) Summary: The Site Reliability Engineer (SRE) is responsible for improving the reliability, availability, performance, and operability of client’ supported software systems. This role combines software engineering and IT operations to automate operational work, monitor system performance, and reduce toil. The SRE establishes and manages monitoring, alerting, incident response, and problem management practices to ensure applications remain available and performant during updates and failures. The role partners with engineering, architecture, and product teams to define reliability standards and production readiness requirements. SRE is a practical implementation of DevOps focused on maintaining software quality in fast-paced development environments. Job Description: Required: Strong grounding in SRE/DevOps practices: incident management, blameless post-mortems, SLOs/SLIs, error budgets, production readiness.Experience building/operating monitoring and alerting and using logs/metrics to diagnose issues.Automation/scripting skills (e.g., Python, PowerShell, Bash) and ability to reduce manual operational work.Strong understanding of cloud-based platforms such as Azure Data Bricks + Unity Catalog, AWS S3 and RDS.Strong experience in ETL / ELT work.Understanding of CI/CD concepts, safe deployment patterns, rollback strategies, and change risk controls. Preferred Experience with cloud environments and infrastructure-as-code.Experience with large datasets (multi-million row datasets).Experience with container orchestration and modern runtime platforms (where applicable).Experience building dashboards and reliability reporting for executives and delivery teams.