Principal Engineer, AI Platform

Bmo Us — United States · Posted ~2 hours ago

Senior Full-time

Skills

AI platform engineering platform services production operations cloud architecture identity policy enforcement observability operability AWS Azure Microsoft AI AI Gateway Policy Engine Identity Fabric AI Registry AI Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Build and operate the core platform infrastructure that enables safe, governed, and scalable enterprise AI. You will own shared platform services spanning cloud environments, identity, policy enforcement, registries, runtime guardrails, and observability, with hands-on production responsibility and on-call participation.

Highlights

Hands-on ownership of enterprise AI platform infrastructure, with responsibility for building and operating shared services across major cloud environments. The role emphasizes scalability, governance, operability, latency, and production ownership.

Description

Job Description Principal Engineer, AI Platform & Fabrics Description BMO is building the platform capabilities that make enterprise AI safe, governed, and scalable. We are seeking experienced Principal/Senior engineers to build and operate the core infrastructure that governs how AI runs at BMO — the AI Gateway, Policy Engine, Identity Fabric, AI Registry, Guardrails Runtime, and AI Observability. This is a build-and-run engineering role. You will be part of the team that own Enterprise AI Platform capabilities end to end: Building, configuring and operating them in production, including on-call. You will not build the AI models or applications themselves (those are domain-owned); you build the governed platform they run on and the runtime evidence that proves they run within policy, across AWS, Azure, and Microsoft AI surfaces, under OSFI and OCC expectations. You are a hands-on engineer who has built shared platform services at scale, cares deeply about operability, latency, and correctness, and understands that in a regulated bank the infrastructure must produce its own evidence. You are energized by taking real engineering assets that includes an existing developer portal, an AI registry, a body of policy-as-code, and gateway integrations and hardening, scaling, enhancing and governing them into enterprise-grade platform capabilities. You raise the technical bar for those around you and mentor as you build. What You'll Build & Operate Depending on your specialization, you will own one or more of the following capability areas: Enterprise Control Plane Portal & Registry — a federated AI Registry (agents, models, tools, channels, evaluations across 16+ asset types) with self-service onboarding and lifecycle workflows; federation with external registries (Agent 365, AgentCore, MLflow).Policy Engine — policy-as-code infrastructure (Cedar/OPA), a policy compilation pipeline, GitOps-based domain-scoped bundle distribution, risk-tiered approval workflows, and a policy simulation sandbox.Observability & Audit — a multi-pipeline telemetry architecture (operational + security + compliance), OpenTelemetry GenAI conventions, cross-pipeline trace correlation, lineage-stamped traces, and a 7-year tamper-evident audit lake producing regulator-ready evidence.Governance & Lifecycle — certification workflows, automated compliance scoring, decommission governance, and evidence generation for architecture and model-risk review. Domain Orchestration Gateway Runtime — domain-hub deployment across AWS and Azure; an inline enforcement engine performing request-time policy evaluation, routing, residency, budget/quota, and circuit breaking within strict tiered latency budgets (Fast