Cloud Site Reliability Engineer

Scotiatech — Colombia · Posted ~23 hours ago

Senior

Skills

Google Cloud Platform Site Reliability Engineering Cloud infrastructure Distributed systems Monitoring Automation System resilience Software engineering Systems engineering SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Help engineer reliable, resilient cloud infrastructure for critical enterprise applications. You will build and operate distributed systems, improve monitoring and automation, reduce operational toil, and partner with development teams to apply SRE principles from design through production.

Highlights

Focus on availability, performance, resilience, and automation for large-scale cloud systems. The role combines software and systems engineering while promoting proactive reliability practices and reducing operational workload.

Description

Purpose The System Reliability Engineering team focuses on System Reliability, Reliability Engineering, and Resilience across Global Corporate Functions, with a specific emphasis on Google Cloud Platform engineering and onboarding to Cloud infrastructure. The Cloud Site Reliability Engineer Specialist is responsible for the availability, performance, and reliability of critical corporate function applications hosted on Google Cloud Platform. This role combines software and systems engineering to build and run large-scale, distributed, fault-tolerant systems. The primary objective is to enhance system resilience and reduce operational work through proactive monitoring, automation, and continuous improvement. The position's mandate is to drive the adoption of SRE principles and practices, collaborating with development teams to engineer scalable and reliable solutions from inception through to production. Accountabilities • Design, build, and maintain scalable and resilient infrastructure on Google Cloud Platform using Terraform. • Develop and manage automation solutions with Python to eliminate operational toil and improve system efficiency. • Deploy, manage, and scale containerized applications using Kubernetes, ensuring optimal performance and availability. • Implement and administer secure cloud networking architectures, including virtual private clouds, firewall rules, and load balancing. • Establish and enforce robust security protocols and secret management practices within the cloud environment. • Drive the adoption of GitOps methodologies using tools like Argo CD or Flux for declarative infrastructure and application management. • Define, track, and report on Service Level Indicators and Service Level Objectives to maintain reliability targets. • Construct and optimize CI/CD pipelines to enable rapid, reliable, and automated software delivery. • Lead incident response efforts for Google Cloud Platform, facilitate blameless post-mortems, and implement corrective actions to prevent future occurrences. • Collaborate with application development teams to integrate reliability best practices into the software development lifecycle. Education / Experience / Other Information • Completion of a post-secondary degree in Computer Science, Engineering, or a related technical field. • Experience working within the financial services or a similarly regulated industry. • Expert-level knowledge of Google Cloud Platform services, cloud networking, and security principles. • Familiarity with regulatory and compliance standards applicable to the banking sector. • Extensive experience with Infrastructure as Code, specifically using Terraform. • Advanced proficiency in scripting and automation using Python. • Demonstrated expertise in container orchestration with Kubernetes. • Strong understanding and practical application of GitOps principles and tools such as Argo CD or Flux. • Proven experience designing, building, and maintaining CI/CD pipelines for automated deployments. • In-depth knowledge of SRE principles, including Service Level Objectives, error budgets, and blameless post-mortems. Working Conditions Work in a standard office-based environment; non-standard hours are a common occurrence