Description
DevOps / SRE Lead
Position Overview
We are looking for a DevOps / SRE Lead with strong hands-on technical expertise and an architectural mindset to own the reliability, architecture evolution, automation, and major incident response of our production platform.
You will define infrastructure and reliability standards, lead complex incident response and architecture decisions, mentor engineers, and collaborate with Engineering, Security, Product, and leadership teams.
This role requires practical expertise across Linux, cloud infrastructure, networking, Kubernetes, databases, middleware, and application delivery.
Key Responsibilities
Own the technical roadmap, architecture planning, operational standards, and continuous improvement of production, pre-production, and test environments.Lead the design of a secure, scalable, observable, and recoverable platform across cloud infrastructure, networking, compute, storage, containers, databases, and middleware.Act as the technical incident commander for major production outages: assess issues quickly, contain impact, coordinate engineering teams and cloud vendors, restore services, lead post-incident reviews, and ensure permanent corrective actions are delivered.Design and implement high-availability and resource-isolation strategies, including production/pre-production separation, critical/non-critical workload isolation, data/compute separation, multi-node and multi-AZ deployment, backup, and disaster recovery.Lead the evolution of Linux, cloud platforms, Kubernetes/EKS, CI/CD, Infrastructure as Code, observability, and security-hardening capabilities.Define production readiness standards for new services, covering capacity, resource limits, high availability, monitoring, alerting, logging, tracing, backup, rollback, and operational ownership.Lead performance and reliability investigations across CPU, memory, disks, IOPS, networking, DNS, Kubernetes scheduling, databases, caches, message queues, and application dependency chains.Build and continuously improve monitoring, alerting, on-call response, incident escalation, capacity forecasting, runbooks, and automated recovery mechanisms.Establish safe release and change-management practices, including CI/CD, canary releases, rollback plans, change reviews, release windows, and risk assessments.Drive Infrastructure as Code, operational automation, and platform standardization to reduce manual risk and single points of failure.Mentor DevOps/SRE engineers and guide development teams on reliability, Kubernetes, cloud operations, production readiness, and incident-response best practices.Partner with Engineering, Security, and leadership teams to prioritize resilience investments and communicate technical risks, architecture decisions, and recovery plans clearly.Requirements
10+ years of experience in DevOps, SRE, infrastructure engineering, cloud platform engineering, or production operations; experience as a technical lead or Tech Lead is preferred.Proven experience owning and continuously improving the reliability of medium- to large-scale production environments.Strong leadership during high-pressure incidents, with the ability to make sound decisions from incomplete information, assign clear ownership, communicate status, and drive service recovery.Strong systems architecture skills, with the ability to understand the end-to-end path from user requests through gateways, services, caches, message queues, databases, and third-party dependencies.Deep Linux troubleshooting expertise across CPU, memory, disks, file systems, I/O, networking, TCP, DNS, processes, and threads.Hands-on experience with AWS or other major cloud platforms, including IAM, VPC, routing, security groups, load balancing, storage, DNS, CDN, WAF, and disaster recovery.Production Kubernetes experience, ideally with EKS or other managed Kubernetes platforms; strong knowledge of node groups, scheduling, networking, storage, Ingress, autoscaling, cluster upgrades, workload isolation, and high availability.Experience operating critical middleware and data platforms, including MySQL/RDS, Redis, Kafka/MSK, Elasticsearch/OpenSearch, Nacos, and XXL-Job.Familiarity with observability tools such as Prometheus, Grafana, ELK/Loki, CloudWatch, OpenTelemetry, or equivalent platforms.Strong Shell scripting skills and the ability to develop automation tools in Python or Go; experience with Terraform, Ansible, Helm, Argo CD, Jenkins, GitLab CI, or similar tools.Ability to define technical standards, produce clear architecture and operational documentation, and drive cross-functional collaboration.Preferred Qualifications
Experience supporting payment, financial services, trading, risk-control, high-availability, or high-concurrency systems.Real experience leading recovery from large-scale node failures, widespread service timeouts, disk I/O saturation, Kubernetes incidents, database failures, or cloud-network incidents.Experience designing multi-AZ, multi-region, active-active, or disaster recovery architectures.Experience with cloud cost optimization, FinOps, capacity planning, and performance engineering.Relevant certifications such as CKA/CKS, AWS Solutions Architect, AWS DevOps Engineer, or equivalent.