Staff Operational Support Engineer

Dolbylaboratories — Australia · Posted ~1 day ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

Role Summary Dolby OptiView is building a dedicated Operational Support (L2) team responsible for the stability, availability, and operational excellence of our 24/7 live video streaming, ads, player, and real‑time delivery platforms. As an Operational Support Engineer (L2), you take end‑to‑end ownership of customer‑impacting production incidents once they are triaged by Level 1 support. You operate directly on production systems, lead live incident resolution, and act as the operational bridge between Support, Engineering, DevOps, and customers, particularly during high‑impact live events. This is a hands‑on, customer‑facing role focused on incident ownership, production operations, automation, and operational scalability, not just reactive troubleshooting. Key Responsibilities Incident & Operational Support Take ownership of escalated customer issues from Level 1 Support and drive them to resolutionTroubleshoot and resolve complex, high-impact production incidents affecting live streams, VOD playback, ad insertion, DRM, and real-time WebRTC servicesOperate directly on production environments, including configuration changes, CDN adjustments, and corrective actions, following established operational procedures, including executing mitigations and emergency changes during live incidents when customer impact requires immediate actionLead or actively contribute to live incident bridges involving customers, internal teams, and partnersProvide clear, timely communication during incidents, including status updates and customer-facing explanations Infrastructure as Code & Production Operations Work fluently with Infrastructure as Code (IaC) to understand, troubleshoot, and safely modify production environmentsLeverage tools and frameworks such as:TerraformHelmKubernetes manifestsGitOps workflowsCI/CD and deployment pipelinesUse IaC as the primary mechanism for safe, auditable, and repeatable operational changesCollaborate with Engineering and DevOps to improve deployment reliability and operational safetyValidate and execute infrastructure or configuration changes through codified workflows AI-Driven Operations & Automation Leverage AI tools and automation to enhance operational efficiency and incident responseContribute to and use: AI-assisted incident triage and classificationAutomated runbook executionAI-based pattern detection across incidentsIntelligent alert correlation and noise reductionUse AI to:Generate or improve incident communicationsAccelerate troubleshooting workflowsIdentify recurring patterns and systemic issuesDrive adoption of automation-first and AI-augmented operational practices Pre-Event Planning & Operational Readiness Participate in pre-event readiness planning for critical customer eventsValidate system readiness through:Runbook checksMonitoring coverage validationRisk identification and mitigation planningDefine and rehearse incident response strategies for high-risk scenariosCollaborate with customers and internal teams to ensure smooth event execution On-Call & 24/7 Operations Participate in a 24/7 on-call rotation, including nights, weekends, and holidays, as part of a global support modelEnsure smooth handovers between shifts and regionsRespond to critical alerts within defined SLAs for stream health, player errors, and delivery infrastructure Root Cause & Continuous Improvement Perform or contribute to root cause analysis (RCA) for production incidentsDocument findings, corrective actions, and preventive measuresIdentify recurring issues and work with Engineering and Product teams to eliminate them permanentlyContribute to and improve runbooks, operational playbooks, and knowledge bases for all OptiView products (Player, ads, live and real time streaming) Collaboration & Engineering Feedback Loop Work closely with Engineering teams to escalate defects, validate fixes, and support production deploymentsProvide feedback on system observability, tooling gaps, and operational risksAct as the operational voice during post-incident reviews Technical Skills Required Skills & Experience 5+ years of relevant experience in operational, support, or similar customer‑facing rolesProven ability to own complex problems end‑to‑end and operate with a high degree of autonomyStrong experience supporting production video streaming platforms, OTT services, live systemsSolid troubleshooting skills across distributed systems (APIs, microservices, cloud infrastructure)Familiarity with HLS, DASH, CMAF, WebRTC, DRM and CDN architecturesExperience working with monitoring, alerting, and logs to diagnose live incidents (Grafana, Kibana/ELK, Prometheus, Loki)Correlate backend streaming metrics, player telemetry, and CDN signals to diagnose live customer issues end-to-end.Comfort performing controlled changes in production environmentsWorking knowledge of incident management and on-call operations Operational Mindset Proven ability to remain calm, structured, and decisive during high-pressure incidentsStrong sense of ownership and accountability for customer outcomesExcellent written and verbal communication skills, including customer-facing communication during incidents