Artificial Intelligence Platform Engineer

Gala Solutions Inc — Canada · Posted ~4 hours ago

Senior Contract Hybrid

Skills

AI engineering platform engineering cloud infrastructure software development enterprise AI AI platforms cloud machine learning infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A large enterprise is seeking a senior AI platform specialist to build scalable AI development platforms, improve adoption, and deliver secure technology solutions across business units.

Highlights

Build enterprise AI platforms and help scale secure, governed artificial intelligence capabilities across an organization.

Description

Job Title: AI Platform Engineering Specialist Experience Level: Level 3 (senior): 5-7 years 12-month contract position Location: Montreal (Day 1 onboarding onsite/in-office presence 3x/week) Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm’s technology principles and drives efficiency and consistency, controls security and strong governance, and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses. This role is for a platform engineering specialist who will help build a firmwide AI development platform and drive adoption of AI capabilities throughout the enterprise. We have multiple focus areas across the platform and are looking for energetic, multi-disciplinary candidates who are eager to contribute to providing scalable, secure, enterprise-wide solutions for the firm. The ideal candidate will have strong hands-on experience building software platforms on any combination of the following platforms—Kubernetes, cloud (AWS, Azure, and/or Google), API-based development, REST framework, data engineering, large-scale API gateway environments, etc. Knowledge of AIML and hands-on experience implementing solutions using generative AI are also preferable. The candidate will have great communication skills, a team-based mentality, and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements. In the Technology division, we leverage innovation to build the connections and capabilities that power our firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. Key responsibilities: Design, build, and operate the AI Gateway's Azure and AWS deployments, taking them from proof of concept to production.Develop and extend the Python services (FastAPI / Flask) that provide the Gateway's inference, onboarding, and administrative APIs.Integrate new model providers and model families, including Azure AI Foundry / Azure OpenAI and AWS Bedrock, covering request signing, streaming responses, failover, and quota handling.Implement cloud-native authentication and secrets handling—Entra ID with Managed Identity and workload federation, AWS IAM roles and STS—with the goal of eliminating stored credentials.Build and evolve the entitlement and authorization data layer across SQL Server and PostgreSQL, including schema changes, migrations, and data-correctness controls.Own the platform controls that make the Gateway a governance point: rate limiting, token accounting, content guardrails, audit logging, and chargeback reporting.Deploy and run the service on Kubernetes (on-premises, AKS, and EKS) using Helm, GitOps, and Terraform, and keep the CI/CD pipelines (Jenkins, GitHub Actions) healthy.Build the observability to answer any question about a request after the fact—metrics, logs, and dashboards across Prometheus, Grafana, Loki, and Snowflake.Work with cloud platform, network, and security teams on connectivity, egress policy, network controls, and architecture review, and produce the evidence those reviews require.Support production: participate in on-call, investigate incidents, and drive fixes and hardening back into the code.Write tests and documentation as part of delivery, and review peers' changes. Required qualifications: Strong, production-grade Python, including a web framework—FastAPI or Flask—and a real testing discipline.Hands-on Kubernetes: deploying, configuring, and troubleshooting workloads, not solely reading manifests.Practical OIDC / OAuth 2.0: token validation, JWKS, client-credentials flows, and claim and audience handling.Microsoft Azure, hands-on across at least three of AKS, Entra ID (app registrations, service principals, Managed Identity / Workload Identity), Azure OpenAI or Azure AI Foundry, Key Vault, Azure Database for PostgreSQL, Azure Cache for Redis, and Azure Monitor.Amazon Web Services, hands-on across at least three of: IAM and STS / assume-role, SigV4 request signing, Bedrock, EKS, VPC endpoints and private networking, Secrets Manager, and CloudWatch.Infrastructure as code—Terraform, Bicep, or CDK—and CI/CD with Jenkins or GitHub Actions.SQL and relational data modeling, including schema migrations.Clear written and verbal communication and the ability to work directly with security, network, and platform teams. Preferred qualifications: Experience building or operating an API gateway, reverse proxy, or multi-tenant platform.LLM platform engineering specifics: streaming and server-sent events, token accounting, prompt and response guardrails, and model evaluation.Kafka and Snowflake for audit and consumption data pipelines.Observability depth: Prometheus and PromQL, Grafana, Loki, OpenTelemetry.Redis or Valkey beyond basic caching—counters, TTLs, distributed rate-limiter semantics.Experience delivering in a regulated enterprise environment with corporate proxies, private networking, and strict change control.