Summary
Join a 24/7 Network Operations Center team responsible for monitoring the health, availability, and performance of enterprise infrastructure, networks, cloud services, applications, databases, and storage. You will investigate alerts, perform first-level troubleshooting, manage incidents, coordinate escalations, maintain operational documentation, and contribute to service reliability and continuous improvement.
Highlights
Work in a structured 24/7 operations environment with exposure to enterprise infrastructure, cloud services, networking, incident management, and continuous improvement. The role offers broad technical collaboration and opportunities to strengthen monitoring, troubleshooting, and operational expertise.
Description
Job Description
The NOC Monitoring Engineer is responsible for continuously monitoring the availability, performance, and health of IT infrastructure, network devices, cloud services, applications, databases, and enterprise systems within a 24ร7 Network Operations Center (NOC).
The role ensures timely identification, analysis, escalation, and coordination of incidents to minimize service disruptions and maintain agreed service levels.
The incumbent performs first-level monitoring and troubleshooting, manages incident tickets, and collaborates with technical support teams to ensure operational excellence and service continuity.
Responsibilities
Detailed Roles and Responsibilities:
Monitor servers, network devices, applications, cloud services, databases, and storage infrastructure using enterprise monitoring tools.
Continuously monitor system availability, performance, capacity, and service health.
Identify, validate, and analyze monitoring alerts and events to distinguish genuine incidents from false positives.
Perform first-level troubleshooting and conduct initial impact assessments.
Create, update, track, and manage incident tickets within the IT Service Management (ITSM) platform.
Escalate incidents to the appropriate resolver groups in accordance with defined service-level agreements (SLAs) and escalation procedures.
Monitor infrastructure metrics including CPU utilization, memory consumption, disk usage, network connectivity, service availability, application response, URL availability, and related performance indicators.
Coordinate with Infrastructure, Network, Cloud, Database, Security, and Application Support teams to facilitate incident resolution.
Participate in major incident management activities, bridge calls, and operational communications when required.
Prepare daily operational status reports, monitoring summaries, shift handover reports, and dashboard updates.
Maintain operational documentation, standard operating procedures (SOPs), runbooks, and knowledge base articles.
Ensure compliance with organizational monitoring standards, operational processes, security policies, and service-level agreements.
Support continuous improvement initiatives related to monitoring, alert management, incident response, and operational efficiency.
Ensure timely communication and escalation of critical incidents to relevant stakeholders.
Work as part of a 24ร7 shift operation and support business continuity requirements.
Qualifications
Educational Qualifications:
Bachelor's Degree in Computer Science, Information Technology, Information Systems, Telecommunications, Engineering, or a related discipline.
Professional certifications in IT Operations, Monitoring, or Infrastructure Management are considered an advantage.
Skills & Experience
Minimum 5+ years of experience in a Network Operations Center (NOC), IT Monitoring, IT Operations, or Technical Support environment.
Experience working with enterprise monitoring platforms such as SolarWinds, Dynatrace, Zabbix, BMC Helix, SiteScope, PRTG, or similar solutions.
Hands-on experience with IT Service Management (ITSM) tools such as ServiceNow, ManageEngine, Jira, or equivalent platforms.
Understanding of Windows and Linux operating systems.
Basic knowledge of networking concepts including TCP/IP, DNS, DHCP, routing, switching, WAN, and LAN technologies.
Exposure to cloud platforms such as Microsoft Azure, Oracle Cloud Infrastructure (OCI), Amazon Web Services (AWS), or Google Cloud Platform (GCP).
Knowledge of virtualization technologies and enterprise infrastructure environments.
Understanding of ITIL Incident Management, Problem Management, Event Management, and Change Management processes.
Experience in alert monitoring, event correlation, incident response, and operational reporting.
Strong analytical, troubleshooting, communication, and documentation skills.
Ability to prioritize and manage multiple incidents in a fast-paced operational environment.
Willingness to work in a 24ร7 shift-based environment.
Behavioral Skills
Strong analytical and problem-solving capability.
High attention to detail and accuracy.
Customer-focused mindset with commitment to service excellence.
Effective verbal and written communication skills.
Strong teamwork and collaboration abilities.
Ability to work under pressure and manage critical incidents professionally.
Accountability and ownership of assigned tasks and activities.
Strong organizational and time-management skills.
Adaptability and willingness to learn new technologies and processes.
Proactive approach to monitoring, incident identification, and continuous improvement.