Platform Engineer

Flexkube — Canada · Posted ~3 hours ago

Skills

Cloud infrastructure Infrastructure as Code Terraform CI/CD Deployment automation Change management Observability Reliability engineering Rollback strategies Production operations

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Platform Engineer to automate software delivery and production operations across cloud infrastructure, data pipelines, servers, and application services. You will design infrastructure as code, build robust CI/CD pipelines, improve observability and reliability, and develop self-service tools that help engineering teams ship changes safely. The role offers the opportunity to influence engineering productivity and operational excellence without participating in an on-call rotation.

Highlights

Build foundational infrastructure and automation that enables engineering teams to deliver software safely and reliably. Work across cloud systems, deployment pipelines, observability, and developer enablement, with no on-call rotation requirement.

Description

We are seeking a Platform Engineer to automate software delivery and production operations across data pipelines, servers, and application services. The role spans infrastructure provisioning, change management, deployment through development, staging, and production, observability, and reliability engineering. You will build the systems and automation that help engineering teams deliver changes safely and on-call responders detect, diagnose, and recover from failures. This role does not participate in the on-call rota. Key Responsibilities • Design and manage cloud infrastructure using infrastructure as code, including Terraform. • Build and maintain CI/CD pipelines that test and promote changes through development, staging, and production, with appropriate approval controls, canary deployments where appropriate, telemetry-based observation periods, and rollback capabilities. • Implement developer platform capabilities, templates, and self-service infrastructure. Build actionable monitoring and alerting, including alert routing and escalation, test that alerts reach the right responders, and measure incident response times using alerting and incident-management data. • Implement OpenTelemetry instrumentation, traces, logs, and metrics to measure service reliability, service response times, and performance against agreed service-level objectives (SLOs) and service-level agreements (SLAs). • Develop and test automation for provisioning and updating environments, virtual machines, and application services. • Build and test diagnostic tools, runbooks, and recovery automation with on-call responders so they can restore service efficiently. • Improve platform security, reliability, scalability, and operational efficiency. • Work with application and data engineers to apply deployment and reliability practices across services and data pipelines. • Use AI-assisted development tools such as Codex to build automation, and review and test the resulting code. Required Skills & Experience • Strong practical experience with infrastructure as code, including Terraform. • Experience building, testing, and operating CI/CD pipelines for controlled promotion across environments, deployment monitoring, and rollback. • Experience operating cloud-native environments. • Practical experience applying Site Reliability Engineering (SRE) principles to production systems, including reliability measurement, actionable alerting, and recovery automation. • Experience with observability platforms, monitoring, and alerting frameworks. • Strong infrastructure automation and scripting skills, including automated testing of infrastructure and deployment changes. • Understanding of platform security, governance, and operational controls. • Experience diagnosing production failures and building tools that help engineering teams and incident responders resolve them. • Ability to understand the purpose and wider context of a requirement, explain assumptions and trade-offs, and judge when to proceed, clarify, or challenge the proposed approach. Preferred Qualifications • Experience with OpenTelemetry. • Experience supporting Databricks and AI/ML platforms. • Familiarity with Kubernetes, container platforms, and cloud automation services. • Experience implementing canary deployments and evaluating telemetry before wider rollout. • Knowledge of developer platform engineering practices. • Experience using AI coding agents in software development.