Site Reliability Engineer

Codeanalytiqa Ltd — United States · Posted ~1 hour ago

Mid Full-time Onsite

Skills

Python observability incident response automation cloud migration platform modernization

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a team responsible for the reliability and operational health of a critical platform. You will monitor and support core infrastructure, lead incident response, build automation tools, improve observability, and work on cloud migration projects while collaborating with product teams to enhance system performance.

Highlights

Work on platform reliability and observability, drive automation initiatives, resolve infrastructure incidents, modernize core systems, and contribute to operational efficiency in global markets.

Description

Job Descriptions: The team owns the reliability and operational health of the Quartz platform. We actively monitor and support core infrastructure and components, execute automation initiatives, build and enhance internal tools, resolve infrastructure issues, and lead incident response to keep the platform running at scale. We are looking for a Site Reliability Engineer (SRE) who will operate hands-on across the stack to improve platform and application observability, drive reliability improvements, and deliver measurable gains in operational efficiency across Global Markets. This role will work closely with Quartz core teams to execute on platform modernization, Quartz migration to Cloud (Cloud 1.0, 2.0), harden production systems, and evolve support tooling. This position is critical to maintaining execution velocity, reducing operational risk, and ensuring Quartz continues to meet its reliability and performance objectives. Job Responsibilities: Strong Python development experienceHands-on Django and REST API developmentStrong MySQL/database skillsDeep Linux administration and troubleshooting experienceInfrastructure engineering and automation backgroundExperience with Ansible and CI/CD pipelinesObservability and monitoring experience (Dynatrace preferred)Experience working in large-scale production environmentsAzure and/or AWS cloud experienceSRE, Reliability Engineering, and DevOps practicesOpenTelemetry and modern observability toolsInfrastructure-as-Code and automation frameworksFinancial Services / Trading Platform experience