Summary
✨ AI‑Generated
A technology organization is looking for a senior platform engineer to operate and improve cloud-native infrastructure. The role covers Kubernetes management, automated deployments, security, and monitoring systems.
Highlights
Work on large-scale infrastructure platforms with responsibility for reliability, automation, and deployment processes.
Description
Senior Platform Engineer
We are seeking a Senior Platform Engineer to join the NaaS Platform team, responsible for the deployment, operation, and reliability of the NaaS platform infrastructure.
You will manage the Kubernetes clusters, GitOps pipelines, infrastructure-as-code, and monitoring/alarming frameworks that underpin a multi-environment telecommunications automation platform.
Responsibilities
Deploy, operate, monitor, and maintain Kubernetes clusters hosting NaaS components, including CI/CD workers, operators, ingest, Kafka, API gateways, collectors, and databases.Design, operate, maintain, and monitor GitOps workflows using Flux for environment consistency and automated deployments.Manage Vault administration for secrets management and the centralised credential lifecycle across all components.Develop and maintain Helm charts and deprecated Ansible playbooks for component deployment and upgrades.Implement and operate CI/CD pipelines using GitLab for build, test, and deployment automation.Integrate with and monitor ancillary telecommunications systems such as Nexus, ISE, LDAP, and DNS.Build out and integrate with cloud-provider solutions.Drive production rollout processes and provide environment support across all environments.Conduct chaos and Litmus testing to validate single-site and multi-site high-availability mitigations.Design, implement, and maintain geo-HA strategies for critical NaaS platform components.Design, implement, and maintain homogeneous telemetry and observability solutions that provide crucial insights into platform health.
Key Skills – Required
Linux systems expertise, including internals, debugging, and performance tuning.Kubernetes production operations, including multi-namespace management, networking, RBAC, resource management, and troubleshooting.GitOps tooling, such as Flux or Argo CD, for declarative infrastructure management.Infrastructure as Code using Terraform, Ansible, and Helm.CI/CD pipeline design and operations using GitLab CI.Vault by HashiCorp, or a similar secrets-management platform.Strong scripting and automation skills using Python and Bash.Multi-cluster Kubernetes management and hybrid-cloud operations.Security by design, including image scanning and hardening.Elasticsearch, Kafka, and InfluxDB cluster management and operations.Prometheus stack, Dynatrace, APM, structured-log ingestion and enrichment, tracing, and OpenTelemetry.BDD smoke testing of deployments and integrations.
Key Skills – Desirable
Container-image build optimisation and registry management.PostgreSQL administration and high-availability configuration.Integration with OpenStack APIs, Cinder, and Nova.Experience with chaos-engineering frameworks such as LitmusChaos or Chaos Mesh.Experience with other orchestration or cloud platforms such as OpenStack.