Site Reliability Engineer
Smartfox Llc — Canada · Posted ~2 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
What you’ll do
Define and track SLIs, SLOs, and error budgets with product and engineering teams, and use them to guide release and reliability decisionsBuild and run highly available, scalable services on AWS, Azure, or GCP with Kubernetes, Docker, and infrastructure-as-code (Terraform, Pulumi)Write software in Python, Go, or similar languages to automate toil away instead of working around itBuild observability that tells the truth: metrics, logs, traces, and alerts using Prometheus, Grafana, Datadog, OpenTelemetry, or ELKLead incident response: detect, mitigate, communicate, and learn through blameless postmortems with real follow-throughDesign for failure: capacity planning, load testing, chaos experiments, and disaster recovery drillsImprove CI/CD pipelines and release safety with progressive delivery, canaries, and automated rollbacksTune performance and cost across infrastructure without sacrificing reliabilityCut alert noise so on-call pages are rare, meaningful, and actionableWrite runbooks, architecture reviews, and production readiness checklists teams actually usePartner with developers, security, and platform teams to build reliability into services from the start
Find your level
Level Typical experience What you’ll ownSRE I - 0-2 years : Monitoring, automation scripts, on-call shadowing, runbooks, and well-scoped reliability projects, learning with a dedicated mentorSRE II - 2-5 years : Owning reliability for services end to end, leading incident response, building tooling, improving SLOs, mentoring newer engineersSenior SRE - 5+ years : Architecting resilient systems, setting reliability standards, leading cross-team initiatives, shaping on-call and incident culturePromotion criteria are written down at every level, with regular career conversations.
You can grow as a deep technical expert or toward leadership.
Both paths are real here.
What you bring
Must have
Solid Linux and networking fundamentals (DNS, TCP/IP, load balancing, TLS)Programming or scripting ability in Python, Go, Bash, or similarHands-on experience with at least one major cloud providerFamiliarity with containers and Kubernetes conceptsA troubleshooting mindset: you follow evidence, stay calm in incidents, and fix root causesClear communication, in writing and on an incident bridgeDegree or diploma in Computer Science, Engineering, IT, or a related field, or equivalent experience.
Home labs, bootcamps, and self-taught engineers are welcome.
We hire for skill, not pedigree.Legal authorization to work in Canada (citizens, permanent residents, and valid work permit holders)
Nice to have (not required)
Experience with Terraform, Ansible, Helm, Argo CD, or GitOps workflowsObservability tooling: Prometheus, Grafana, Datadog, Splunk, OpenTelemetryKnowledge of distributed systems, databases, and caching layers (PostgreSQL, Redis, Kafka)Certifications: CKA, AWS Solutions Architect or DevOps Professional, Azure AZ-400, Google Professional Cloud DevOps EngineerExperience with incident management tools (PagerDuty, Opsgenie) and postmortem practiceChaos engineering, performance testing, or capacity planning experienceFamiliarity with security and compliance practices (SOC 2, PIPEDA)Industry background in [fintech, SaaS, e-commerce, healthcare, telecom, gaming, etc.]Fluency in French, an asset for roles serving Quebec or bilingual teams [if applicable]A home lab, GitHub repos, or open-source contributions you can talk about
We have 154,634 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 154,634 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume