Mid-to-Senior Cloud SRE & DevOps Engineer

Mondrian Alpha — United States · Posted ~6 hours ago

Senior

Skills

Cloud engineering Site Reliability Engineering DevOps Kubernetes Public cloud GitOps Infrastructure automation Observability Security Production engineering Cloud SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A mid-to-senior cloud infrastructure role for an engineer who takes ownership of production systems. You will design and operate resilient environments across public cloud and Kubernetes, enable development teams through modern DevOps and GitOps practices, and improve reliability, observability, automation, security, and operational excellence across critical platforms.

Highlights

Mid-to-senior opportunity centered on production ownership across cloud, Kubernetes, SRE, DevOps, and platform engineering. The role emphasizes reliability, observability, automation, security, GitOps, and operational excellence while supporting business-critical technology platforms.

Description

Mid-to-Senior Cloud SRE & DevOps Engineer Overview A Global Asset Manager is looking for a mid-to-senior level engineer to join a Cloud SRE & DevOps Engineering team focused on building, operating, and evolving AI, cloud infrastructure, and software delivery platforms that support trading, investment, and other business-critical systems. This role sits at the intersection of AI Infrastructure Engineering, Cloud Engineering, Site Reliability Engineering (SRE), DevOps, Platform Engineering, and Production Engineering. It is designed for individuals who take ownership of systems running in production. You will be responsible for designing and operating resilient, scalable environments across public cloud and Kubernetes platforms while enabling engineering teams through modern DevOps and GitOps practices. The position reflects a strong SRE mentality, with an emphasis on reliability, observability, automation, security, and operational excellence. You will contribute to cloud transformation initiatives, ensuring systems are built for performance, stability, scalability, and resilience. This is a hands-on role requiring deep technical expertise, accountability for production systems, and a mindset oriented toward continuous improvement, risk management, and engineering efficiency. The role also contributes to evolving platform capabilities supporting AI, machine learning, and data-intensive workloads. You will work closely with software engineering, cybersecurity, data, and quantitative technology teams to deliver secure, scalable, and high-performing systems while improving developer experience and platform maturity across the organization. What We're Looking For Core Technical Expertise Strong hands-on experience with public-cloud platforms in production environments.Experience working in hybrid or multi-cloud environments.Deep experience with Kubernetes and containerized systems.Strong familiarity with Kubernetes deployment and management tooling.Proven experience with infrastructure as code, preferably Terraform or a comparable technology.Strong experience building and managing CI/CD pipelines.Experience with Git-based and GitOps workflows.Proficiency in Python, Shell, PowerShell, or similar scripting languages for automation.Production & Systems Engineering Strong understanding of distributed systems, networking, and cloud architecture.Experience operating and supporting production systems with high availability and performance requirements.Experience diagnosing and resolving complex production incidents in distributed environments.Experience supporting relational databases and cloud-based data platforms.Understanding of backup, disaster recovery, resiliency, and business continuity practices.Strong troubleshooting and problem-solving skills.Monitoring & Observability Hands-on experience with enterprise monitoring and observability platforms or comparable technologies.Experience designing metrics, alerts, dashboards, and operational monitoring for production systems.Strong understanding of logging, tracing, alerting, and proactive system health monitoring.Experience improving system observability and reducing operational risk.