Site Reliability Engineer - New Grad

Jobright Ai — Canada · Posted ~2 hours ago

Junior Full-time

Skills

Site reliability engineering Monitoring Service-level indicators Service-level objectives Automation Incident response Scalability Observability SRE AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an engineering team as a New Grad Site Reliability Engineer and help keep an AI-powered platform reliable, scalable, and observable. You will define and monitor service-level objectives, build dashboards and alerts, automate operational tasks, investigate incidents and performance issues, and improve system resilience and recovery. The role offers high ownership and hands-on exposure to production AI infrastructure.

Highlights

An entry-level opportunity to work on production AI systems with high ownership and direct impact. The role provides hands-on experience in reliability, automation, observability, incident response, scalability, and resilient system design.

Description

Jobright is your personal AI job search agent that transforms the way you do job search from solo, time-consuming efforts to a fast, expert-guided journey, simplifying every job search step and accelerating your route to the best job outcomes. The New Grad Site Reliability Engineer will help keep the AI-powered platform reliable, scalable, and observable by automating operations and improving production reliability. Why Join Us • Build real, production AI agents used by real users • High ownership and impact • Work at the intersection of AI, agents, and product • Shape how people experience AI-driven job search Responsibilities • Help define and monitor service-level indicators and objectives • Build dashboards, alerts, and reliability reporting • Automate operational tasks and reduce manual intervention • Investigate incidents, performance issues, and system failures • Improve service scalability, resilience, and recovery procedures • Partner with engineers on production readiness and safe releases • Document runbooks and contribute to incident reviews Qualification Required • Recent graduate in Computer Science, Engineering, Information Systems, or a related field • Programming or scripting experience with Python, Go, Bash, or a similar language • Familiarity with Linux, networking, databases, and distributed-system fundamentals • Strong analytical and debugging skills • Ability to communicate clearly in a remote environment Preferred • Experience with cloud platforms, containers, or Kubernetes • Familiarity with metrics, logs, traces, and observability tools • Understanding of SRE concepts such as SLIs, SLOs, and error budgets • Previous internship or systems-focused project experience