GCP Site Reliability Engineer

Tekgence Inc — Canada · Posted ~5 hours ago

Senior Full-time

Skills

Google Cloud Platform GCP infrastructure Terraform cloud monitoring incident management CI/CD cloud security Python Bash high availability GCP Compute Engine GKE Cloud Storage VPC IAM Cloud SQL Load Balancing Cloud Monitoring Cloud Logging PowerShell

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A GCP Site Reliability Engineer will design, deploy, and maintain cloud infrastructure while improving reliability, automation, monitoring, and production support. You will manage core cloud services, implement infrastructure as code, build CI/CD pipelines, troubleshoot incidents, perform root-cause analysis, and strengthen security, backup, disaster recovery, and high availability.

Highlights

Infrastructure-focused SRE role working with modern GCP services, infrastructure as code, automation, monitoring, production reliability, security, and high-availability systems.

Description

Job Summary We are looking for a GCP Infrastructure & Cloud SRE Engineer to support and maintain cloud infrastructure for a banking client. The role will focus on Google Cloud Platform (GCP), infrastructure automation, monitoring, reliability, and production support. Key Responsibilities Design, deploy, and maintain infrastructure on Google Cloud Platform (GCP).Manage GCP services including Compute Engine, GKE, Cloud Storage, VPC, IAM, Cloud SQL, and Load Balancing.Implement Infrastructure as Code using Terraform.Monitor system performance, availability, and reliability using Cloud Monitoring and Cloud Logging.Support production incidents, troubleshoot infrastructure issues, and perform root-cause analysis.Develop and maintain CI/CD pipelines for cloud infrastructure and applications.Implement automation using Python, Bash, or PowerShell.Manage cloud security, IAM roles, service accounts, and network configurations.Support disaster recovery, backup, and high-availability solutions.Work with development and application teams to improve system reliability and performance.Participate in on-call support and incident management activities.Ensure infrastructure follows banking security, compliance, and governance standards.Required Skills 5+ years of experience in Cloud Infrastructure / SRE / DevOps.Strong hands-on experience with GCP.Experience with Terraform and Infrastructure as Code.Strong knowledge of GKE/Kubernetes, networking, IAM, and cloud security.Experience with CI/CD tools such as Jenkins, GitLab, or Azure DevOps.Experience with monitoring and logging tools.Strong Linux/Unix administration and scripting skills.Experience with Python or Bash scripting.Strong troubleshooting and production support experience.