Summary
✨ AI‑Generated
Join an enterprise operations environment responsible for monitoring infrastructure, applications, databases, networks, and cloud services. You will triage alerts, manage incidents and escalations, support outage communications, monitor dashboards, and contribute to SLA reporting using established IT service management practices.
Highlights
Broad IT operations role covering enterprise monitoring, incident response, infrastructure visibility, outage communication, and service management, with exposure to major enterprise monitoring platforms.
Description
Must Have Skills
The candidate should have strong experience in Enterprise Command Center Operations, Event Monitoring, Incident Management, Major Incident Management, Service Desk Operations, and IT Operations Support.
The individual should possess hands-on experience with monitoring tools such as Splunk, SolarWinds, Nagios, Zabbix, Dynatrace, AppDynamics, SCOM, OEM, or similar enterprise monitoring platforms.
Strong understanding of ITIL processes, including Incident Management, Problem Management, Change Management, and Service Request Management, is essential.
The candidate should be proficient in monitoring infrastructure, servers, databases, network devices, cloud services, and business-critical applications.
Experience in alert triaging, escalation management, ticket handling, dashboard monitoring, outage communication, and SLA reporting is mandatory.
Strong analytical, troubleshooting, communication, and stakeholder management skills are required.
Nice To Have Skills
Experience with ServiceNow, BMC Remedy, Opsgenie, PagerDuty, Microsoft Teams integrations, Power Automate, Splunk Dashboard development, Cloud Monitoring (AWS, Azure, OCI), Automation, Python, PowerShell, SQL, Grafana, Elasticsearch, Kibana, Ansible, and DevOps tools would be advantageous.
Exposure to Site Reliability Engineering (SRE), AIOps platforms, workflow automation, and reporting tools such as Power BI is also beneficial.
Detailed Job Description
The Command Center Engineer will continuously monitor infrastructure, applications, databases, networks, and cloud services using enterprise monitoring tools.
The role is responsible for identifying alerts, validating incidents, performing initial troubleshooting, raising tickets, and coordinating with resolver groups for timely resolution.
The engineer will manage incident lifecycles, support major incident bridges, communicate service impacts to stakeholders, and ensure compliance with SLA and operational procedures.
The individual will perform event correlation, alert analysis, impact assessment, outage tracking, escalation management, and service restoration activities.
The role includes monitoring dashboard health, reviewing recurring incidents, identifying service trends, generating operational reports, and contributing to process improvement initiatives.
The Command Center Engineer will maintain operational documentation, runbooks, knowledge articles, escalation matrices, and monitoring procedures.
The role also involves participating in shifts, handling priority incidents, coordinating change activities, validating scheduled maintenance windows, and supporting business continuity and disaster recovery exercises.