Terraform / GCP DevOps Engineer

Dabstergroupuk — Sweden · Posted ~9 hours ago

Senior

Skills

Google Cloud Platform Terraform Infrastructure as Code CI/CD Cloud infrastructure Platform reliability Cloud security Disaster recovery Business continuity GitHub Actions Grafana GCP Cloud Run Cloud Functions

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced DevOps engineering role focused on designing and operating resilient cloud infrastructure. You will architect production environments on a major cloud platform, manage infrastructure as code, optimize CI/CD, improve reliability and security, and lead disaster recovery and continuity practices.

Highlights

Experienced cloud engineering role focused on scalable GCP infrastructure, infrastructure as code, CI/CD, reliability, security, cost efficiency, disaster recovery, and mentoring.

Description

JOB DESCRIPTION TERRAFORM / GCP DEVOPS ENGINEER (Cloud Infrastructure) POSITION OVERVIEW We are seeking an experienced Terraform/GCP DevOps Engineer to join our managed engineering team delivering infrastructure and cloud operations You will be responsible for: Designing, building, and maintaining enterprise-grade infrastructure on Google Cloud Platform (GCP)Managing Infrastructure-as-Code (Terraform) across production environmentsBuilding and optimizing CI/CD pipelines for continuous deploymentEnsuring platform reliability, security, and cost efficiencyLeading disaster recovery, backup strategies, and business continuityMentoring other engineers on infrastructure and cloud best practices KEY RESPONSIBILITIES Infrastructure Design & Architecture Design and architect scalable, resilient cloud infrastructure on GCP for production workloadsDefine infrastructure patterns, standards, and best practices for the platformEvaluate cloud services and tooling (Cloud Run, Cloud Functions, Firestore, BigQuery, Pub/Sub, etc.)Plan capacity and auto-scaling strategies to meet SLA requirements (99%+ availability)Design disaster recovery (DR) and business continuity (BC) architectures (RTO: 1 day, RPO: 24 hours)Document architecture decisions (ADRs) and maintain up-to-date system design documentationConduct architecture reviews with engineering teams to ensure alignment with IKEA standardsInfrastructure-as-Code & Terraform Build and maintain Terraform modules for all GCP resources (compute, networking, storage, databases, security)Manage Terraform state securely with remote state backend (GCS + Terraform Cloud)Implement infrastructure versioning, testing, and code review workflowsAutomate infrastructure provisioning, scaling, and updatesEnsure infrastructure code is modular, reusable, and well-documentedMaintain 100% Infrastructure-as-Code coverage (zero manual changes in production)Plan and execute infrastructure updates with zero-downtime deploymentsImplement drift detection and auto-remediation for infrastructure state CI/CD Pipeline & Deployment Automation Design and maintain GitHub Actions CI/CD pipelines for infrastructure and application deploymentsImplement automated testing for infrastructure code (Terraform validation, policy checks, security scanning)Build deployment automation and orchestration (Blue-green deployments, canary releases)Manage secrets and credentials securely (Google Secret Manager, encrypted Terraform variables)Implement automated rollback procedures and chaos engineering testsMonitor deployment quality metrics (change failure rate, mean time to recovery)Work with Cloudflare for edge deployment and traffic management Cloud Operations & Reliability Manage GCP project structure, IAM policies, and security controlsImplement monitoring, observability, and alerting across all infrastructure componentsSet up and optimize cloud resource monitoring (Cloud Monitoring, Cloud Logging, Trace)Manage cloud budgets, cost optimization, and resource efficiencyExecute routine operational tasks: backups, maintenance windows, security patchingParticipate in on-call rotation for infrastructure incidents (P1/P2)Troubleshoot infrastructure issues and optimize performanceMaintain detailed runbooks and disaster recovery procedures Security, Compliance & Cost Management Implement infrastructure security best practices (network segmentation, encryption, IAM least privilege)Manage firewall rules, VPC configuration, Cloud Armor, and DDoS protectionEnsure compliance with security standards, data residency, and privacy requirementsConduct infrastructure security audits and vulnerability assessmentsManage data backup and encryption strategiesImplement cost optimization strategies (reserved instances, committed use discounts, resource optimization)Monitor and reduce cloud spend while maintaining performanceManage cloud provider relationships and licensing Mentoring & Knowledge Transfer Mentor junior/mid-level engineers on infrastructure, cloud platforms, and DevOps best practicesLead knowledge transfer sessions on Terraform, GCP, and CI/CD best practicesContribute to team documentation and internal knowledge baseParticipate in code reviews for infrastructure and deployment changesLead design reviews and architecture discussions REQUIRED QUALIFICATIONS Experience 5+ years of professional DevOps/Cloud Infrastructure engineering experience3+ years of hands-on experience with Terraform in production environments (preferably managing 50+ resources)3+ years of production experience on Google Cloud Platform (GCP) or equivalent (AWS/Azure) with deep expertise in 2+ major GCP services2+ years managing CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or similar)Proven experience building and maintaining disaster recovery (DR) and business continuity (BC) strategiesExperience operating customer-facing, high-availability services with 99%+ uptime targetsTrack record of incident response and on-call operations in production environments Technical Skills - GCP & Cloud Platforms Core GCP Services Compute: Deep expertise with Cloud Run (containerized workloads), Cloud Functions (serverless), Compute EngineDatabases: Firestore (NoSQL), BigQuery (data warehouse), Cloud SQL (relational databases)Messaging & Streaming: Pub/Sub, DataflowStorage: Cloud Storage (GCS), Firestore Backup & Point-in-Time Recovery (PITR)Networking: VPC, Cloud Load Balancer, Cloud Armor, VPN, Cloud InterconnectSecurity: Cloud IAM, Secret Manager, Cloud KMS, Security Command CenterMonitoring & Observability: Cloud Monitoring, Cloud Logging, Cloud Trace, ProfilerDeployment & Orchestration: Cloud Deploy, GKE (optional but valuable) Infrastructure-as-Code & Terraform Advanced Terraform proficiency: modules, variables, outputs, state management, remote backendsTerraform best practices: version control, testing, CI/CD integrationTerraform Cloud / Enterprise (optional, valuable for remote state management)Terraform testing frameworks: Terratest, checkov, terraform validate, tflintPolicy-as-Code: OPA/Conftest or similar for infrastructure policy enforcementHands-on experience managing 50+ cloud resources with TerraformGit workflow and code review processes for infrastructure code CI/CD & Deployment Automation GitHub Actions expertise (workflow design, custom actions, secrets management)GitHub Advanced Security: code scanning, dependency scanning, secret scanningContainer technologies: Docker, container registries (Artifact Registry, Container Registry)Blue-green and canary deployment strategiesInfrastructure-as-Code testing and validation in pipelinesAutomated rollback and disaster recovery proceduresMonitoring deployment quality metrics Observability & Monitoring OpenTelemetry instrumentation and integrationGoogle Cloud Monitoring and Cloud LoggingSentry for error tracking and alertingDashboard design and visualization (Grafana, Cloud Monitoring dashboards)Alert design, incident triage, and escalation proceduresMetrics collection, tracing, and debugging distributed systems Networking & Security VPC design, subnetting, routing, NAT/PATCloud Load Balancer and traffic managementCloud Armor (DDoS protection, WAF)VPN and encrypted communicationIAM policy design and least-privilege accessSecret management and credential rotationNetwork security auditing and compliance Disaster Recovery & Backup Backup strategy design (RTO/RPO planning)Firestore backup and point-in-time recoveryGCS versioning and retention policiesInfrastructure disaster recovery testingRunbook development and documentation PREFERRED QUALIFICATIONS Experience with Cloudflare (Workers, edge computing, WAF/DDoS)Kubernetes/GKE experience (container orchestration, Helm, operators)Experience with data pipelines (Dataflow, Cloud Composer/Airflow)Cost optimization expertise (RI analysis, committed use discounts, workload-based sizing)Experience with multi-region or multi-cloud deploymentsBackground in platform engineering or building internal developer platforms (IDP)Exposure to governance and compliance (SOC 2, ISO 27001, GDPR)Experience with chaos engineering and resilience testing (Gremlin, Chaos Toolkit)Knowledge of API Gateway, Apigee or similarExperience with service mesh (Istio, Consul) - optionalTerraform Cloud or Terraform Enterprise administrationLinux/Unix system administration background