Site Reliability Engineer

Infotree Global Solutions โ€” Canada ยท Posted ~6 days ago

Mid Remote

Skills

SRE Observability AWS Google Cloud Azure Terraform Kubernetes Datadog Sentry Golang Python Node.js

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Help improve reliability and observability for a cloud platform by building monitoring solutions, managing infrastructure, and supporting large-scale distributed systems.

Highlights

Fully remote role focused on observability, monitoring, platform reliability, and cloud infrastructure.

Description

Description: We are looking for a Site Reliability/ Observability Engineer to help ensure that our Product and Platform Engineers can monitor and observe our platform while continuing to rapidly ship software that our customers love. If you have experience within the Site Reliability Engineering (SRE) field or working as a Development Operations (DevOps) engineer, and you have a passion for Observability tooling, this position will allow you to further your learning and development in these areas. We are looking for engineers who are passionate about monitoring, observing, measuring uptime and availability, and ensuring stability for our platform. Skills & Qualifications: 3+ years of platform operations engineering, SRE, or DevOps experience.Experience with cloud infrastructure like AWS, Google Cloud, or Azure.Experience with Datadog (preferred) or other monitoring tools.Experience with Sentry (preferred) or other error reporting tools.Experience managing infrastructure with Terraform.Proficiency in Golang, Node.js, or Python.Experience with Kubernetes.Demonstrable expertise in monitoring distributed applications at scale.Understanding of microservice architecture and best practices.