Site Reliability Engineer

Selby Jennings — United States · Posted ~3 hours ago

Senior Hybrid Visa History ✓

Skills

Site reliability engineering Python Automation Linux administration High Performance Computing Infrastructure engineering Production troubleshooting Linux HPC Distributed computing

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an infrastructure engineering team responsible for the reliability and performance of critical high-performance systems. You will build Python-based automation, operate Linux environments and HPC clusters, troubleshoot production incidents, improve observability, and help scale distributed compute platforms.

Highlights

Hybrid SRE role focused on reliability, scalability, automation, and performance of critical infrastructure, with hands-on exposure to HPC clusters, distributed computing, observability, and production systems.

Description

Site Reliability Engineer Location: Chicago, IL (Hybrid) A leading proprietary trading firm is seeking a Site Reliability Engineer to support and enhance the reliability, performance, and scalability of critical trading infrastructure. Responsibilities Build and maintain reliable, high-performance infrastructure supporting trading and research systems.Develop automation and operational tooling using Python.Monitor system health, troubleshoot production issues, and drive performance improvements.Support and optimize HPC clusters and distributed compute environments.Improve observability, incident response, and operational processes.Collaborate with software engineers to enhance platform reliability and efficiency.Requirements 4+ years of experience in Site Reliability Engineering, Production Engineering, or Infrastructure Engineering.Strong Python programming and automation skills.Experience supporting High Performance Computing (HPC) environments.Strong Linux administration and troubleshooting experience.Excellent problem-solving skills in mission-critical environments.Preferred Networking knowledge (TCP/IP, routing, switching, network troubleshooting).Experience supporting low-latency or high-throughput systems.Background in proprietary trading, quantitative finance, or other performance-sensitive environments. This is an opportunity to work on mission-critical systems at the heart of a high-performing trading organization, alongside talented engineers in a fast-paced and highly technical environment. This is a hybrid role based out of the firms Chicago office requiring 3 days of onsite work per week, and a rotational on call schedule.