Site Reliability Engineer

Objectivepartners — United States · Posted ~2 hours ago

Senior Full-time

Skills

cloud infrastructure automation monitoring ETL high performance computing Cloud Monitoring Automation HPC

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A reliability engineering position focused on building automated infrastructure, data pipelines, monitoring systems, and scalable compute environments for demanding applications.

Highlights

High-impact infrastructure role focused on automation, reliability, and scalable systems. Opportunity to work with advanced computing environments and engineering teams.

Description

Job Title: Site Reliability Engineer (Infrastructure & Systems) Location: Chicago, IL (Greater Metro Area) About The Opportunity Join a premier financial technology organization at the intersection of high-performance computing, low-latency execution systems, and scalable modern infrastructure. We are seeking a proactive Site Reliability Engineer to architect, automate, and elevate our real-time trading ecosystem while working directly alongside quantitative researchers and core engineering teams. This is a high-impact role designed for a technical problem solver who thrives in fast-paced, high-concurrency environments. Responsibilities Infrastructure Automation: Engineer self-healing, automated platform solutions to maximize operational reliability across mission-critical execution environments. Data & Compute Pipelines: Construct and maintain data ingestion workflows (ETL) alongside high-performance cloud compute (HPC) environments. System Telemetry: Design customizable monitoring, alerting, and reporting suites to deliver proactive visibility into system health and network performance. API & Service Integration: Build and integrate robust RESTful APIs connecting multi-cloud platforms with existing on-premises data systems. Database & Network Support: Oversee local and cloud-hosted SQL/NoSQL datastores while collaborating with core engineering to automate network workflows. Incident Escalation: Serve as a Tier-2 escalation point to rapidly diagnose, troubleshoot, and resolve system anomalies during active operations. Security & Best Practices: Maintain enterprise-grade operational standards, ensuring all deployed infrastructure complies with firm-wide security protocols. Requirements (Must-Have) 3+ years of professional experience in Site Reliability Engineering, Systems Administration, DevOps, or Network Engineering. Strong programming proficiency in Python and Unix shell scripting, paired with Git version control. Demonstrated experience with Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation, Ansible, or Puppet). Practical understanding of containerization frameworks (such as Kubernetes) in hybrid or public cloud settings. Familiarity with building and maintaining continuous integration and deployment (CI/CD) pipelines (e.g., GitHub Actions, GitLab CI, Jenkins). Strong foundational knowledge of TCP/IP networking concepts (DNS, Routing, VPNs) and enterprise storage solutions (NFS, block, object). Excellent collaborative communication skills with a track record of thorough technical documentation. Preferred Qualifications (Nice-to-Have) Hands-on expertise with major cloud platforms (AWS or Google Cloud Platform / GCP). Familiarity with managed container orchestrators such as Amazon EKS or Google GKE. Exposure to low-latency financial systems, quantitative research tools, or high-performance computing environments. Compensation & Benefits Competitive base salary paired with performance-based incentive opportunities. Comprehensive benefits package including medical, dental, vision, 401(k) retirement options, and paid time off. Equal Opportunity Employer: All qualified candidates will receive consideration without regard to race, color, religion, sex, national origin, protected veteran status, or disability status.