Senior IT Engineer - Infrastructure Operations

Sotalentjobs — United States · Posted ~2 hours ago

Senior Full-time Hybrid $116800-$175200 annually

Skills

IT infrastructure cloud technologies incident management monitoring networking systems operations cloud platforms monitoring tools observability platforms

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A large organization is seeking a senior infrastructure engineer to support enterprise systems, improve reliability, and handle critical technology operations using modern cloud and monitoring solutions.

Highlights

Senior infrastructure role focused on operational resilience, cloud technologies, and solving complex enterprise technology challenges.

Description

Senior IT Engineer – Infrastructure Operations Location: Charlotte, NC Work Arrangement: Hybrid – 3 days per week in office - This is a hybrid role based in Charlotte, NC OR Hartford, CT, with an expectation of working from the office three days per week, Tuesday through Thursday. Industry: Insurance / Financial Services Employment Type: Full-time Salary: $116,800–$175,200 annually A leading organization in the insurance and financial services sector is seeking a Senior IT Engineer – Infrastructure Operations to support enterprise technology operations and major incident management. This role focuses on using AI-driven analysis, cloud technologies, monitoring and observability platforms to rapidly identify, assess and resolve critical service issues. The position will work closely with infrastructure, application, reliability engineering, service desk and vendor teams to improve operational resilience and service restoration. Key Responsibilities Manage technical and executive communications during major incidents.Perform alert triage, event correlation and initial impact assessments.Identify actionable events and distinguish them from non-actionable monitoring noise.Support major incident detection, escalation and service restoration.Investigate and prioritize events using established SOPs, runbooks and decision frameworks.Maintain situational awareness during high-impact technology events.Document event patterns, operational observations and recurring issues.Promote events to incidents when appropriate and coordinate with resolver teams.Identify opportunities to improve monitoring, alert quality, automation and operational readiness.Collaborate with infrastructure, application, reliability engineering, service desk and external vendor teams.Recommend improvements to monitoring, runbooks, workflows and automation based on recurring operational issues.Required Qualifications 8+ years of experience working with monitoring and observability technologies such as Splunk, Dynatrace, ITSI, Moogsoft, ThousandEyes or similar platforms.Experience using AI-driven data analysis and problem-solving techniques.Proficiency with AI tools, including prompt engineering and investigative/interrogation techniques.Familiarity with ITSM and ticketing platforms, including incident creation, categorization, escalation and documentation.Understanding of event correlation, alert prioritization and service-impact analysis.Experience with basic automation and workflow enablement.Strong analytical and problem-solving abilities.Excellent operational judgment and the ability to make decisions under pressure during high-impact incidents. Additional Details The position may require shift coverage, after-hours support and escalation responsibilities as part of a 24x7 operational model and follow-the-sun support structure. Compensation: $116,800–$175,200 base salary, with total compensation potentially including additional incentive and benefit components. Work authorization: Candidates must be authorized to work in the United States without employer sponsorship.