Summary
✨ AI‑Generated
A technology-focused organization is seeking a DevOps Engineer to support critical cloud platforms, improve operational reliability, automate diagnostics, and troubleshoot distributed systems across production environments.
Highlights
Opportunity to work on large-scale cloud infrastructure, production operations, automation, observability, and reliability engineering while solving complex technical challenges.
Description
Position: DevOps Engineer
• Provide L3 operational support for Mesh-hosted APIs, applications and shared platform services across production and non-production environments.
• Lead structured diagnosis and root cause analysis for outages, degraded services, failed deployments, routing faults, certificate issues, connectivity problems and platform alerts.
• Operate and troubleshoot Azure Kubernetes Service workloads, including deployments, pods, services, ingress, daemonsets, jobs, configuration, secrets, probes, resource constraints and cluster events.
• Support Kong Gateway and Consul Connect, including ingress, service discovery, service routing, service resolvers, service routers and service-to-service connectivity.
• Use logs, metrics and traces from Splunk, Elastic/Kibana, OpenTelemetry and application monitoring tools to isolate issues and validate recovery.
• Build read-only diagnostics, health checks and operational automation using Bash, Python, Ansible, kubectl and jq.
• Support platform releases, blue/green migrations, restoration activities, environment validation, change implementation and post-change verification.
• Troubleshoot CI/CD and deployment workflows involving Bamboo, Bitbucket, Git, Helm, Maven and enterprise artifact repositories.
• Support integrations with Kafka/Confluent Cloud, Redis, Cassandra/Datastax Astra, Vault, certificates and external API-management or service-management systems.
• Work with application, cloud, network, security and service-management teams to restore service, communicate impact and drive preventive actions.
• Maintain runbooks and knowledge articles, improve support tooling, and reduce repetitive manual effort through practical automation.
• Participate in rostered on-call support and planned after-hours or weekend changes when required.