AI Platform Engineering Specialist

Randstadenterprise β€” Canada Β· Posted ~6 hours ago

Hybrid

Skills

AI platform engineering AWS Microsoft Azure Python FastAPI Flask Cloud-native development API development Authentication and authorization SQL Server PostgreSQL Infrastructure operations Azure Azure AI Foundry Azure OpenAI AWS Bedrock Entra ID AWS IAM

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AI platform engineering role based in Montreal with initial office presence required three days per week. You will take cloud-based AI gateway deployments from proof of concept to production, develop Python APIs, integrate foundation-model providers, implement cloud-native identity and secrets management, and build authorization and data-layer controls across relational databases.

Highlights

Build and operate production AI gateway infrastructure across Azure and AWS. The role combines Python API development, cloud security, model-provider integration, data-layer engineering, and platform governance, with significant ownership of production AI infrastructure.

Description

Job Title: AI Platform Engineering Specialist Location: Montreal (Day 1 onboarding onsite/in office presence 3x/week) Key responsibilities Design, build and operate the AI Gateway's Azure and AWS deployments, taking them from proof of concept to production.Develop and extend the Python services (FastAPI / Flask) that provide the Gateway's inference, onboarding and administrative APIs. Integrate new model providers and model families, including Azure AI Foundry / Azure OpenAI and AWS Bedrock, covering request signing, streaming responses, failover and quota handling. Implement cloud-native authentication and secrets handling β€” Entra ID with Managed Identity and workload federation, AWS IAM roles and STS β€” with the goal of eliminating stored credentials. Build and evolve the entitlement and authorization data layer across SQL Server and PostgreSQL, including schema changes, migrations and data-correctness controls. Own the platform controls that make the Gateway a governance point: rate limiting, token accounting, content guardrails, audit logging and chargeback reporting. Deploy and run the service on Kubernetes (on-premises, AKS and EKS) using Helm, GitOps and Terraform, and keep the CI/CD pipelines (Jenkins, GitHub Actions) healthy. Build the observability to answer any question about a request after the fact β€” metrics, logs and dashboards across Prometheus, Grafana, Loki and Snowflake. Work with cloud platform, network and security teams on connectivity, egress policy, network controls and architecture review, and produce the evidence those reviews require. Support production: participate in on-call, investigate incidents, and drive fixes and hardening back into the code.Write tests and documentation as part of delivery, and review peers' changes. Required qualifications: Strong, production-grade Python, including a web framework β€” FastAPI or Flask β€” and a real testing discipline.Hands-on Kubernetes: deploying, configuring and troubleshooting workloads, not solely reading manifests.Practical OIDC / OAuth 2.0: token validation, JWKS, client-credentials flows, claim and audience handling. Microsoft Azure, hands-on across at least three of: AKS, Entra ID (app registrations, service principals, Managed Identity / Workload Identity), Azure OpenAI or Azure AI Foundry, Key Vault, Azure Database for PostgreSQL, Azure Cache for Redis, Azure Monitor. Amazon Web Services, hands-on across at least three of: IAM and STS / assume-role, SigV4 request signing, Bedrock, EKS, VPC endpoints and private networking, Secrets Manager, CloudWatch. Infrastructure as code β€” Terraform, Bicep or CDK β€” and CI/CD with Jenkins or GitHub Actions.SQL and relational data modelling, including schema migrations. Clear written and verbal communication, and the ability to work directly with security, network and platform teams. Preferred qualifications Experience building or operating an API gateway, reverse proxy or multi-tenant platform. LLM platform engineering specifics: streaming and server-sent events, token accounting, prompt and response guardrails, model evaluation. Kafka and Snowflake for audit and consumption data pipelines. Observability depth: Prometheus and PromQL, Grafana, Loki, OpenTelemetry. Redis or Valkey beyond basic caching β€” counters, TTLs, distributed rate-limiter semantics. Experience delivering in a regulated enterprise environment with corporate proxies, private networking and strict change control.