Site Reliability Engineer, New Graduate

Jobright Ai — Canada · Posted ~18 hours ago

Junior

Skills

Site reliability engineering Monitoring SLIs SLOs Dashboards Alerting Automation Incident investigation Scalability Resilience SRE Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Start your reliability engineering career by helping operate a scalable, observable AI-powered platform. You will define and monitor service-level indicators and objectives, create dashboards and alerts, automate operational work, investigate incidents and performance problems, and strengthen system resilience and recovery. The role provides close collaboration with software engineers and hands-on exposure to production operations.

Highlights

New-graduate opportunity focused on production reliability, scalability, and observability for an AI-powered platform. Provides hands-on exposure to monitoring, automation, incident investigation, resilience, recovery, and production readiness while working closely with software engineers.

Description

Jobright is your personal AI job search agent that transforms the way you do job search from solo, time-consuming efforts to a fast, expert-guided journey, simplifying every job search step and accelerating your route to the best job outcomes. The New Grad Site Reliability Engineer will help keep the AI-powered platform reliable, scalable, and observable by automating operations and improving production reliability. Why Join Us • Build real, production AI agents used by real users • High ownership and impact • Work at the intersection of AI, agents, and product • Shape how people experience AI-driven job search Responsibilities • Help define and monitor service-level indicators and objectives • Build dashboards, alerts, and reliability reporting • Automate operational tasks and reduce manual intervention • Investigate incidents, performance issues, and system failures • Improve service scalability, resilience, and recovery procedures • Partner with engineers on production readiness and safe releases • Document runbooks and contribute to incident reviews Qualification Required • Recent graduate in Computer Science, Engineering, Information Systems, or a related field • Programming or scripting experience with Python, Go, Bash, or a similar language • Familiarity with Linux, networking, databases, and distributed-system fundamentals • Strong analytical and debugging skills • Ability to communicate clearly in a remote environment Preferred • Experience with cloud platforms, containers, or Kubernetes • Familiarity with metrics, logs, traces, and observability tools • Understanding of SRE concepts such as SLIs, SLOs, and error budgets • Previous internship or systems-focused project experience