Summary
✨ AI‑Generated
Own the day-to-day operation, stability, governance, and continuous improvement of enterprise testing, quality engineering, and AI platforms. You’ll manage service performance and risks, lead major incident and root-cause activities, engage stakeholders, and ensure platforms remain reliable, secure, compliant, and fit for business needs.
Highlights
Broad operational ownership across testing, quality engineering, and AI platforms, with a strong focus on reliability, governance, stakeholder engagement, continuous improvement, and service performance.
Description
Job Description – Run Engineer
Position Summary
The Run Engineer is responsible for the operational management, stability, support, governance, and continuous improvement of enterprise testing tools, Quality Engineering Platforms (QEP), and AI platforms.
The role acts as the primary point of accountability for BAU operations, service performance, stakeholder engagement, risk management, and platform lifecycle activities, ensuring services remain reliable, secure, compliant, and aligned with business objectives.
Roles & Responsibilities
Manage the day-to-day operations and support of Testing Tools, QEP, and AI platforms.Ensure platform availability, reliability, performance, and operational health while meeting agreed SLAs.Act as the primary escalation point for critical and major incidents, operational risks, and service disruptions.Lead incident, problem, and root cause analysis activities and ensure permanent fixes and preventive measures are implemented.Manage operational backlogs, Jira boards, ServiceNow queues, and support priorities across multiple teams.Establish and maintain operational governance, support processes, runbooks, and service standards.Drive platform modernization, automation, and DevOps adoption.Work with engineering teams to improve deployment pipelines, release processes, and operational readiness.Review and optimize CI/CD pipelines to improve deployment reliability, speed, and quality.Support infrastructure automation using Chef, Ansible, Terraform, GitHub Actions, Jenkins, Azure DevOps, Bamboo, and OpenShift Pipelines.Develop and maintain automation scripts and operational tools using Python, Shell Scripting, PowerShell, and API-based integrations.Manage platform upgrades, infrastructure changes, patching, and environment lifecycle activities.Collaborate with Cloud, Infrastructure, Network, Security, and Architecture teams to deliver reliable platform services.Monitor platform health, observability dashboards, logging solutions, and proactively identify issues.Drive continuous improvement by identifying recurring issues and automating manual operational tasks.Ensure operational compliance with enterprise security, risk, audit, and governance requirements.Lead vendor engagement, support escalations, contract reviews, and service performance discussions.Provide executive-level reporting on platform availability, incidents, risks, capacity, and service improvements.Coach and mentor Run Engineers and promote operational excellence, DevOps maturity, and engineering best practices.Support strategic platform initiatives, migrations, upgrades, cloud transformation, and technology adoption programs.Key Skills & Experience
Strong knowledge of CI/CD practices and DevOps operating models.Hands-on experience with tools such as Jenkins, GitHub Actions (GHA), Azure DevOps (ADO), or Bamboo.Experience with Chef, Ansible, and Terraform.Strong scripting and automation skills using Java, Python, Shell, or similar languages.Knowledge of OpenShift, Kubernetes, Containers, and Microservices.Experience with Infrastructure Monitoring and Observability Platforms.Strong experience in Incident, Problem, Change, and Release Management.Working knowledge of Jira, ServiceNow, and Confluence.Success Measures
High platform availability and reliability.Strong SLA and service performance achievement.Reduced incident volume and faster service restoration.Successful delivery of upgrades, releases, and service improvements.Positive stakeholder satisfaction and audit outcomes.Increased automation, operational efficiency, and platform adoption.