Platform Engineer

Workaxle β€” Canada Β· Posted ~3 hours ago

Senior Full-time

Skills

Kubernetes Terraform Helm GitOps Cloud Infrastructure Argo

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A platform engineering role focused on building reliable infrastructure across production environments. You will manage Kubernetes platforms, automation workflows, migrations, and shared engineering systems.

Highlights

Own critical platform infrastructure, architecture decisions, and reliability improvements across production environments.

Description

Responsibilities Architecture and Design Design cross-cutting platform solutions spanning multiple services, and record the decisions and trade-offs Design for failure: bounded retries, backpressure, circuit breaking, honest health checks, tested failover β€” and contain blast radius to one tenant, region or subsystem Plan upgrades and migrations that complete without user-facing downtime Establish shared defaults, libraries and patterns that other engineering teams adopt Kubernetes and Infrastructure Own the Kubernetes platform across multiple production regions β€” cluster architecture, upgrades, add-ons, RBAC and network policy Develop and maintain infrastructure-as-code using Terraform, Helm and GitOps workflows (Flux, Argo) Own the shared components everything depends on β€” databases, caches, message brokers, workflow engine β€” including upgrades, patching, replica topology, capacity and disruption budgets Execute storage, instance class and engine version migrations with rollback plans Reliability, Monitoring and Incident Response Lead investigation and resolution of complex production issues spanning services, infrastructure layers and technology stacks β€” eliminating the class of problem rather than the instance Establish monitoring and alerting so that failures are detected before customers report them Define service level indicators and objectives with error budgets, measured at the user-visible boundary Keep alerting focused on user-visible symptoms rather than resource noise, and the paging tier small and trustworthy Participate in the on-call rotation, and convert post-incident reviews into work that gets completed Security and Compliance Maintain automated scanning and posture monitoring across code, dependencies, infrastructure and cloud accounts, with remediation targets Integrate security into the development lifecycle, including threat modelling on new designs Support SOC 2, ISO 27001 and ISO 27018 evidence and audit requirements Required qualifications 8+ years of experience building and operating production software, still working hands-on in code Experience with multi-tenant SaaS at scale β€” real customer load, contractual service levels, and meaningful consequences when systems fail. Experience limited to small or early-stage products is not sufficient Strong system design skills: anticipating edge cases and failure modes before implementation, and judging what will be costly to change later Deep Kubernetes expertise β€” understanding how the platform works internally, not only how to deploy to it. Cluster architecture, upgrades, add-ons, RBAC, troubleshooting Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux, Argo), with an infrastructure-as-code approach in preference to manual configuration Strong hands-on expertise in AWS or an equivalent cloud provider (EC2, S3, RDS, VPC, IAM) Proficiency in at least one backend language (Go, Java, Python, Ruby, TypeScript, C#). We run Ruby and TypeScript, but we hire for fundamentals rather than a language match Desirable qualifications Distributed systems patterns: transactional outbox, idempotency, workflow orchestration, message brokers and their delivery guarantees, service mesh behaviour Reliability practice: service level objectives, error budgets, incident response PostgreSQL at depth β€” query plans, replication, performance under load Monitoring and observability tooling (Prometheus, Grafana, Loki, Datadog or equivalent) CI/CD pipeline design and implementation (GitHub Actions, CircleCI or equivalent) Security and compliance frameworks (SOC 2, ISO 27001, ISO 27018) Experience in a domain where errors are expensive and service levels are strict β€” payroll, financial services, healthcare, telecommunications or similar