Site Reliability Engineer

Remitian — United States · Posted ~6 hours ago

Mid Full-time

Skills

Cloud Infrastructure Infrastructure Automation Observability Monitoring Deployment Pipelines Production Operations Reliability Engineering Automation Cloud CI/CD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A high-agency Site Reliability Engineering role responsible for keeping business-critical cloud systems dependable. You will own meaningful parts of infrastructure, observability, automation, deployment pipelines, and production operations, while participating in a monitoring rotation and continuously engineering manual operational work out of the system.

Highlights

High-ownership SRE role responsible for infrastructure, observability, automation, deployment pipelines, and production reliability in a business-critical environment. The position combines building and operating systems, with direct influence over engineering practices and opportunities to automate manual operational work.

Description

About the Role Remitian is looking for a high-agency Site Reliability Engineer to own the reliability of the platform that moves money and tax data for our customers. Our systems run on a daily clock: payment files are prepared, submitted to banking partners, and reconciled against receipts on a fixed schedule every business day. When something in that chain is late or wrong, customers and tax authorities feel it the same day. You will own the infrastructure, observability, and automation that keep those flows dependable, and you will take your turn in our production monitoring rotation. That rotation is deliberately part of the job: we want an engineer who lives with the operational reality of the platform and then engineers the manual work out of it. This is a build-and-operate role, not a ticket queue. You will own meaningful parts of our cloud footprint, deployment pipelines, and monitoring, and you will have a direct hand in shaping how Remitian runs software in production. This role is based in either our Ottawa, Ontario office or in Florida, with Miami preferred, on a hybrid schedule of 3 to 4 days per week in office. Key Responsibilities Infrastructure and Cloud Platform Own and evolve our cloud infrastructure, defined as code and deployed repeatedly across environmentsBuild and maintain CI/CD pipelines that get changes to production safely and oftenManage environment configuration, secrets, and access in a way that is auditable and least-privilege by defaultImprove deployment mechanics so releases are routine, low-drama, and reversibleManage networking, connectivity, and secure remote access for engineers and servicesReliability and Observability Own monitoring, dashboards, logging, and alerting for the scheduled jobs and services that move money and dataMake alerts actionable: reduce noise, add the context an on-call engineer actually needs, and retire alarms that no longer earn their placeDefine and track service level objectives for the workflows customers depend onLead incident response, drive issues to resolution, and turn postmortems into completed follow-up workImprove error handling, retries, and recovery paths for time-sensitive batch and integration workflowsProduction Monitoring Rotation and Automation Take your turn in the production monitoring rotation, verifying that daily payment and reconciliation workflows complete as expected and investigating when they do notTriage production alerts, diagnose root causes across services, jobs, queues, and third-party integrations, and communicate status clearly to engineering, support, and compliance stakeholdersTreat every manual check as a defect to be automated: replace human verification with automated validation, self-healing jobs, and tooling that surfaces problems before a person has to look for themBuild and improve internal tooling and dashboards that make the rotation faster, safer, and eventually unnecessaryHelp train and onboard other engineers into the rotation, and keep the runbooks currentSecurity and Compliance Support Implement and maintain access controls, audit trails, and infrastructure guardrails appropriate to a regulated fintechKeep dependencies, images, and infrastructure patched and currentSupport audit and compliance requirements with evidence, documentation, and tooling rather than manual effortHandle production data with care, applying the principle of least privilege to yourself as well as to everyone elseOwnership and Continuous Improvement Take real ownership of infrastructure, tooling, and operational workflowsIdentify problems, propose solutions, and drive them to completion without waiting for directionPush back, raise concerns, and surface ideas that improve reliability or how the team worksContribute to engineering standards, technical documentation, and onboardingWho You Are Three to seven years of experience in site reliability, platform, or infrastructure engineeringStrong hands-on experience with a major cloud provider, ideally AWS, including compute, networking, storage, and identityComfortable with infrastructure as code and with CI/CD pipelines as a first-class part of your workComfortable writing code, not just configuration, with scripting and automation in Python, Bash, C#, TypeScript, or similarPractical experience with observability tooling, including logs, metrics, dashboards, and alertingExperience being on call for production systems, and a track record of reducing operational load rather than absorbing itStrong debugging instincts across distributed systems, scheduled jobs, queues, and third-party integrationsHigh agency, meaning you find work that needs doing, take ownership, and move things forward without being toldStrong written and verbal communicator, able to explain incidents, trade-offs, and risk clearly to technical and non-technical audiencesSound judgment around production access, sensitive data, and change managementExcited to work in a fast-paced startup environment where priorities shift and ambiguity is part of the jobBased in or willing to relocate to either the Ottawa, Ontario area or Florida, with Miami preferred, and excited to be in the office 3 to 4 days per weekNice to Haves Experience in fintech, payments, banking, or another regulated industryExposure to batch payment file processing, bank integrations, or financial reconciliation workflowsExperience operating .NET workloads in productionExperience with containers and managed container orchestration servicesExperience with zero-trust or modern remote access toolingExperience building internal tools, dashboards, or developer toolingExperience using LLMs and AI tooling to automate operational work or accelerate developmentExperience in a startup where you wore multiple hats and owned production end to endWhy Join Us Own the reliability of systems that move real money on a daily deadline, where your work is visible and matters immediatelyJoin a high-performing team building mission-critical fintech infrastructureA clear mandate to automate, since we are actively investing in replacing manual operational work with tooling, and you will lead that effortBroad scope across cloud infrastructure, deployment, observability, security, and internal toolingUse modern tools, including AI and LLM tooling, to move quickly and build with leverageHybrid work from either our Ottawa or Florida location, where the team comes together 3 to 4 days per week to build, ship, and learn alongside one another