Summary
✨ AI‑Generated
A cloud operations engineering role responsible for improving reliability, monitoring, automation, and incident management for critical services. The ideal candidate has experience with DevOps practices, cloud platforms, and operational improvement initiatives.
Highlights
Work on critical production systems with a focus on reliability, automation, and operational excellence. The role involves modern cloud operations, observability improvements, and collaboration with engineering teams.
Description
About the Company
Client: UHG
About the Role
Job Title: Azure Systems / Devops Engineer
Work Environment: 24x7 Production Support
Need only Local Candidate
Location: Minneapolis, MN
$60/hr on C2C all inc
Responsibilities
Own reliability outcomes for critical services: Define/track SLIs & SLOs, manage error budgets, and drive continuous improvement to availability, latency, and resiliency.Operate and mature the observability/AIOps platform: Build and tune monitoring, dashboards, alerting, and correlation (logs/metrics/traces) to reduce noise and speed detection and diagnosis.Lead incident response & problem management: Run on-call/war rooms, perform root-cause analysis, publish post-incident reviews, and ensure corrective/preventive actions are delivered.Automate toil and enable self-healing: Create runbooks, scripts, and workflows for automated remediation, safe changes, and guardrails; improve MTTR through automation.Partner with engineering and stakeholders: Consult on architecture/release readiness, capacity planning, and operational standards; communicate reliability risk and KPI trends to leadership.
Qualifications
Prior Experience & Domain Expertise Summary5+ years in SRE/production operations3+ years in observability dashboards building3+ years in Cloud Engineering1+ year of Agentic SRE tooling – build & operationalize
Required Skills
AIOps & SRE fundamentals: 5+ years in SRE/production operations; apply SLO/SLI, error budgets, incident management, and automated remediation patterns.Observability toolset: 3+ years building dashboards/alerts and troubleshooting with tools such as Dynatrace and Splunk (log/metric/trace correlation, alert tuning, noise reduction).Cloud & platform engineering: 3+ years operating services on AWS/Azure/GCP; strong Linux, networking, containers/Kubernetes, and IaC (e.g., Terraform/CloudFormation) for reliable, repeatable environments.Automation & agentic ops mindset: Proficient in Python and/or Bash plus CI/CD; build runbooks, self-healing workflows, and safe change automation; strong troubleshooting, ownership, and collaboration in on-call rotations.AI-enabled SRE / intelligent ops: 1–2+ years applying AI-assisted incident response (auto summarization, auto triage, pattern detection) and predictive alerting/anomaly detection in production ops.Security/DevSecOps collaboration: Working knowledge of vulnerability management, secrets/cert governance, and secure CI/CD gates; ability to partner with cyber teams and vendors
Preferred Skills
Preferred Skills and AttributesRelease safety & resiliency engineering: Experience with progressive delivery (canary/blue-green), chaos testing/DR drills, and performance/capacity engineering; familiarity with well-architected reviews (Azure WARA/Azure Advisor or equivalent).ITSM & on-call tooling integration: Hands-on integrating monitoring signals to ServiceNow, PagerDuty (or equivalent), and building low-friction escalation/war-room workflows.Portfolio communication & stakeholder management: Able to translate reliability risks into exec-ready updates (KPIs like MTTD/MTTR, error budget burn, top recurring toil) and drive cross-team alignment.
Pay range and compensation package
$60/hr on C2C all inc
Equal Opportunity Statement
All contractor resources are expected to demonstrate baseline proficiency in enterprise-approved AI tools as part of their day-to-day responsibilities.
This includes, but is not limited to:
Consistent Use: Maintain a minimum of 90% weekly usage of AI tools such as GitHub Copilot, Microsoft 365 Copilot, and other GenAI platforms approved by the enterprise.Applied Productivity: Leverage AI tools to enhance coding, documentation, data analysis, and decision-making workflows.Continuous Learning: Stay current with evolving AI capabilities and features, and apply them to improve delivery quality and velocity.