Summary
✨ AI‑Generated
Join a globally distributed infrastructure team responsible for scaling and operating mission-critical, low-latency systems under demanding workloads. You will improve performance and reliability across infrastructure layers, automate operations, strengthen observability, handle incidents, and drive long-term preventive improvements.
Highlights
A mission-critical infrastructure role in a globally distributed environment, focused on reliability, performance, security, automation, observability, and self-service infrastructure for high-volume trading workloads.
Description
We're looking for a senior DevOps Engineer to join a globally distributed team to help scale and operate infrastructure that powers millions of trades daily across CeFi and DeFi venues.
You'll play a mission ciritcal role in ensuring performance, reliability and security of the systems in a demanding environment.
Key responsibilities:
Ensure performance and reliability of mission-critical trading infrastructure by proactively identifying and resolving bottlenecks at all layers: compute, network, storage and applicationSupport, operate, and continuously improve highly available, low-latency systems under global trading load, with a strong focus on automation and enabling self-service for internal teams.Participate in a day-time rotational on-call schedule, handling real-time operational incidents, root cause analysis, and long-term preventive improvements.Use metrics and observability tools (Prometheus, Grafana, etc.) to detect anomalies, monitor performance trends, and support incident response.Build, manage, and scale infrastructure-as-code using Terraform and configuration management tools like Ansible.Collaborate closely with trading, engineering, and security teams to align infrastructure improvements with application demands.Implement robust security practices, including IAM, threat detection, and secure CI/CD pipelines, to support high-trust production environments.Apply low-level optimization strategies to improve throughput, reduce latency, and increase system efficiency across diverse environments.
Requirements:
5+ years in a DevOps/SRE roles focussed on high-performance, production-grade systemsProficiency in Linux admin and experience managing AWS cloud infrastructureProficiency with containerised workloads (Kubernetes)Strong scripting and automation capabilities using Python and Bash; knowledge of Go, JavaScript, or TypeScript is a desireableHands-on experience with observability stacks (Prometheus, Grafana) and production incident responseExperience with automation and IaC tools; Ansible, TerraformUnderstanding of performance engineering at both the system and network levels, including tuning, profiling, and capacity planningKnowledge of cloud networking, security best practices, IAM and SIEM toolsExperience working with crypto or finance is desireable