GPON Systems Developer

Maziv β€” South Africa Β· Posted ~18 hours ago

Mid Full-time

Skills

Python network automation OSS API development network monitoring AWS APIs

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Looking for a systems developer to automate network management, monitoring, provisioning, and operational workflows in a high-availability infrastructure environment.

Highlights

Build automation solutions for large-scale network operations with strong focus on reliability and production systems.

Description

Description Role Purpose The GPON Systems Developer works with the Active Network and OSS teams to automate the management and monitoring of Vumatel's GPON access network. The role replaces manual, console-driven network operations with tested, repeatable software: automated configuration and provisioning against Huawei, Calix and Nokia OLT platforms, automated collection and interpretation of alarms, optical performance data and telemetry, and automated auditing of live network state against OSS inventory. The fibre network runs 24/7 and carries live customer services, so every automation this role builds must be safe to run against production, observable while it runs, and reversible when it goes wrong. Reporting Line and team structure Reports to: Principal OSS Solutions Architect. Works alongside: Systems Engineers, Software Engineers and QA Engineers within the OSS Engineering team, together with the Platform/DevOps and Site Reliability functions who own the shared AWS and Kubernetes environments. The role works in close and continuous partnership with the Active Network team, who own the OLT, ONT and element management estate that this role automates against. Key Responsibilities Automating network configuration and management Write, test and review C# services and Python tooling that configure GPON equipment programmatically β€” creating and amending ONT service and line profiles, VLAN and traffic profiles, bandwidth profiles and port configuration β€” across Huawei, Calix and Nokia OLT platforms.Build and maintain the abstraction layer that presents a consistent internal interface over the three vendors' differing northbound interfaces and command syntaxes, so that OSS systems issue one instruction regardless of which vendor's equipment terminates the service.Implement automation against the appropriate interface for each platform β€” vendor northbound APIs, NETCONF/YANG, SNMP, TL1, TR-069/ACS or scripted CLI where no programmatic interface exists β€” and handle vendor-specific error responses, session limits, throttling and timeout behaviour.Build guard rails into every automation that changes live network state: pre-change validation, blast-radius limits, dry-run and diff modes, staged rollout, automatic abort on error thresholds and a tested rollback path.Develop and run automated bulk change campaigns β€” profile migrations, configuration standardisation, firmware upgrade rollouts β€” including the batching, scheduling, progress reporting and per-device verification that each campaign requires.Maintain the credential and access model used by automation to reach network devices, ensuring secrets are held in a managed store, rotated, scoped to the minimum required privilege and excluded from logs. Automating network monitoring, assurance and fault finding Build and maintain the collectors that ingest alarms, syslog, SNMP traps, telemetry and performance counters from the OLT estate, normalise them across vendors into a common event model, and publish them to the platforms used by the NOC.Develop alarm correlation, deduplication and suppression logic so that a single physical event β€” a fibre break, an OLT card failure, a power outage β€” raises one actionable alarm rather than thousands of individual ONT alarms, and so that dying-gasp and mass-outage patterns are recognised automatically.Build automated optical performance monitoring: scheduled collection of ONT and OLT transmit and receive levels, trending of signal degradation over time, and proactive flagging of drops, splices and ONTs deteriorating towards failure before the customer reports a fault.Develop automated diagnostic and line-test tooling that the NOC and Service Desk run against an individual service β€” retrieving ONT status, optical levels, ranging state, uptime, error counters and recent alarm history β€” and return a plain-language result rather than raw device output.Build rogue ONT, unauthorised device and PON instability detection, and the reporting that routes each finding to the correct team for action.Develop automated reconciliation between OSS inventory and live network state, identifying configuration drift, orphaned or stranded service configurations, ONTs present on the network but absent from inventory, and mismatched profiles β€” and produce the exception reports that drive their correction.Monitor the automation and collection platforms themselves for exceptions and failed transactions, investigate the underlying fault, and drive each to resolution β€” whether by correcting data, replaying the operation, patching the integration or escalating to the responsible system or vendor β€” working the outstanding exception backlog down on a scheduled basis rather than allowing it to accumulate. Dat, reporting and persistence Design, review and evolve PostgreSQL schemas holding network state, event history and performance data, including indexing strategy, partitioning of high-volume time-series and event tables, and archival or downsampling of historical records.Write and tune SQL, review execution plans for slow or lock-contending queries surfaced by monitoring and remediate them before they affect collection throughput or dashboard responsiveness.Author and apply versioned database migrations, verifying each against production-like data volumes in a lower environment before release.Build the operational dashboards and scheduled reports used by the Active Network, NOC and Capacity teams β€” PON and OLT utilisation, alarm volumes and top talkers, optical degradation trends, firmware version distribution, automation success and failure rates. Deployment, infrastructure and scalability Containerise services using Docker and maintain their Kubernetes deployment configuration, including resource requests and limits, readiness and liveness probes, horizontal pod autoscaling and rolling update strategies that do not drop in-flight device sessions or lose collected events.Design collection and automation for the scale of the estate: connection pooling and session reuse against OLTs, rate limiting to protect network management interfaces from being overwhelmed, queue-based work distribution and back-pressure handling during alarm storms.Perform load and soak testing ahead of significant releases or footprint growth, document the results, and adjust scaling, batching and polling configuration accordingly.Review collection lag, event throughput, queue depth, device error rates and automation failure trends weekly, and raise remediation work into the team backlog before thresholds are breached. Reliability, monitoring and incident response Build and maintain dashboards, metrics, structured logging and alert rules for every service the incumbent owns, ensuring each production failure mode β€” a stalled collector, an unreachable OLT, a failing automation campaign β€” produces an actionable alert routed to the correct responder.Respond to production alerts within the agreed response times and lead technical triage where automation or monitoring failures affect network visibility or service delivery, including after hours.Support the NOC and Active Network teams during major network incidents by extracting and interpreting device and event data at speed, and by building ad hoc queries or scripts needed to scope an outage.Conduct and document post-incident reviews, produce written root cause analyses, and implement the resulting corrective actions to completion.Write and keep current the runbooks, escalation paths and troubleshooting guides used by the NOC and by other engineers on call. Testing, quality and change control Write unit, integration and contract tests for all new and amended code, and maintain the vendor simulators, lab OLT environments and recorded device fixtures that allow automation to be tested without touching production equipment.Validate every new automation in the lab against the relevant Huawei, Calix and Nokia platforms, and against each firmware version in production use, before it is permitted to run against the live network.Review colleagues' pull requests, with attention to device safety, failure handling, backwards compatibility across vendor and firmware versions, observability and rollback safety.Prepare and submit change records for production releases and for automated change campaigns, define and test the rollback plan for each, and execute changes during approved maintenance windows, which are frequently scheduled outside of business hours to protect customer-facing availability Collaboration, documentation and knowledge sharing Work with the OSS Solutions Architect and Active Network engineers to refine automation designs, and produce interface specifications, sequence diagrams and architecture decision records.Track vendor firmware releases, deprecations and northbound API changes, assess their impact on existing automation, and schedule the remediation work required.Attend daily stand-ups, sprint planning, refinement and retrospective sessions, and provide realistic estimates and status on committed work.Mentor junior and intermediate engineers through pairing, code review and structured handover of on-call and GPON domain knowledge.Maintain accurate technical documentation in the team's knowledge base and record all work against the appropriate tickets in the team's tracking system. Requirements Minimum Requirements (essential) Experience At least five (5) years' experience in a comparable systems, software or network automation development role.Demonstrable hands-on experience with GPON access network technology, including OLT and ONT operation, service and line profiles, PON port and splitter architecture, ranging, and optical power levels and their interpretation.Working experience with Huawei, Calix and Nokia GPON platforms and their associated element and network management systems. Depth on one vendor with practical exposure to the others will be considered, provided the candidate can demonstrate the ability to work across all three.Demonstrable experience automating configuration, provisioning or monitoring against live network equipment, including the safety and verification practices that make such automation acceptable in production.Demonstrable experience operating high-availability systems that run 24/7, including participation in a formal on-call rota, production incident response and root cause analysis.Experience designing for scale and resilience: retry and back-off strategies, idempotency, rate limiting, graceful degradation and recovery from partial failure. Technical C# and the .NET ecosystem, to a level of independently designing, building and supporting production services.PostgreSQL, including schema design, indexing, query tuning and migration management.Network management protocols and interfaces: SNMP, syslog, NETCONF/YANG, TR-069, TL1 or vendor northbound APIs, and device CLI automation.Sound understanding of REST APIs, message queueing and asynchronous processing patterns.Practical experience with monitoring, alerting, structured logging and time-series or event data.Version control (Git), code review practice and CI/CD pipelines.Automated testing across unit, integration and contract levels. Other Strong written communication in English, sufficient to produce runbooks, incident reports and interface specifications used by other teams.Willingness and ability to work in the Bryanston office on a full-time basis.Willingness to participate in a standby rota and to perform scheduled maintenance work outside of normal business hours. Advantageous Skills And Experience Python and Ruby for automation, tooling, data remediation and supporting legacy services.MSSQL, particularly where integrating with or migrating from existing MSSQL-backed business systems.Elasticsearch, for centralised log and event aggregation, search and operational analytics.Background in telecommunications or a telco-adjacent industry, including exposure to OSS/BSS platforms, fault and performance management, service assurance workflows or TM Forum Open APIs.Experience with additional access technologies or vendor platforms β€” XGS-PON, Adtran, ZTE, DZS β€” and with an ACS platform for TR-069 device management.Experience with RADIUS and AAA platforms, including subscriber authentication, authorisation and accounting, and their integration with access network and OSS systems.Experience with time-series databases, streaming telemetry, or observability tooling such as Prometheus, Grafana or the Elastic Stack.Infrastructure-as-code experience (for example Terraform), and experience with event streaming or brokered messaging platforms (for example Kafka or RabbitMQ).Understanding of GIS or physical fibre plant data and its relationship to logical network state.Vendor certification on Huawei, Calix or Nokia access platforms.Experience mentoring engineers or acting as a technical lead on a delivery workstream. Qualifications Essential: A completed Bachelor's degree or a recognised diploma (NQF Level 6 or above) in Computer Science, Software Engineering, Information Systems, Electrical, Electronic or Computer Engineering, Telecommunications, or a related technical discipline.Relevant industry or vendor certifications β€” for example AWS Solutions Architect or Developer Associate, Certified Kubernetes Administrator, or Huawei, Calix or Nokia access network certification β€” are advantageous but not a substitute for the above. Work Level Mid-Level Job Type Permanent Salary Market Related EE Position No Location JHB North