Cloud & AI Engineer

Randstadenterprise โ€” Canada ยท Posted ~3 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

This is a global, multi-discipline team responsible for architecting and delivering secure, robust, and innovative cloud and AI enablement solutions which enable development teams to build and deploy new applications, modernize existing workloads, and safely adopt emerging AI capabilities across the public cloud. Location: Montreal (Day 1 onboarding onsite/in office presence 3x/week) Responsibilities: The Cloud & AI Engineer will be part of the Cloud Business Enablement squad and will be responsible for designing, building, and maintaining enterprise-scale multi-cloud infrastructure across Azure and AWS, while enabling cloud-native AI solutions and agentic AI platforms. The role requires a strong understanding of Landing Zone architecture, cloud security controls, enterprise networking, Kubernetes, Terraform, automation pipelines, LLM fundamentals, prompt engineering, and AI agent development. Key Responsibilities: Cloud Platform and Multi-Cloud Infrastructure Engineering Design, build, and maintain Azure and AWS Landing Zones aligned with enterprise security, compliance, and governance standards.Engineer secure, scalable, and resilient multi-cloud foundations across Azure, AWS, and hybrid on-premises environments.Design and manage shared platform services including networking, identity, security, governance, observability, and infrastructure automation.Architect and implement Hub-and-Spoke, Transit Network, and shared services patterns with appropriate segmentation and security controls.Enable on-premises to cloud connectivity using enterprise patterns such as VPN, ExpressRoute, Direct Connect, Transit Gateway, private connectivity, DNS, routing, and firewall controls.Support cloud migrations, modernization, and workload onboarding activities across application, data, and AI use cases.Implement highly available, resilient, and multi-region architectures across cloud accounts, subscriptions, and environments.Troubleshoot complex infrastructure, networking, identity, security, and platform integration issues across hybrid and multi-cloud environments.DevOps, Automation and Kubernetes Develop, maintain, and enhance reusable Terraform modules and cloud infrastructure blueprints.Use GitHub, GitHub Actions, and CI/CD pipelines to automate infrastructure provisioning, policy enforcement, and operational workflows.Implement and manage Kubernetes platforms including Azure Kubernetes Service (AKS) and Amazon Elastic Kubernetes Service (EKS).Support container-based platforms and cloud-native application workloads with strong focus on scalability, resiliency, observability, and security.Develop automation scripts, libraries, and integration tooling using Python and related technologies.Apply software engineering best practices including source control, code reviews, test automation, secure development, and operational readiness.Azure and AWS Cloud Services Deploy, secure, and operate Azure services including Azure API Management, App Services, Function Apps, Logic Apps, Azure SQL, Azure Databricks, Key Vault, Storage Accounts, Azure Monitor, Log Analytics, Entra ID, App Registrations, Subscriptions, Management Groups, Azure Policy, and tagging governance.Deploy, secure, and operate AWS services including EC2, S3, RDS, Lambda, VPC, DynamoDB, CloudFront, IAM, Transit Gateway, CloudWatch, and related cloud-native services.Integrate cloud environments with third-party enterprise tools such as Splunk, Akamai, F5, Datadog, CloudWatch, Azure Monitor, and related logging, monitoring, load balancing, WAF, CDN, and observability platforms.AI Engineering and Agent Development Design, develop, and deploy enterprise-grade AI agents and autonomous workflows using modern AI frameworks and cloud-native architectures.Apply LLM fundamentals including transformer concepts, tokenization, embeddings, inference patterns, context windows, grounding, retrieval patterns, and model limitations.Apply prompt engineering techniques including prompt design, prompt optimization, prompt chaining, structured outputs, system instructions, guardrails, and reusable prompt patterns.Implement AI agent state management including memory handling, context persistence, session management, orchestration, and workflow state tracking.Develop and maintain AI Agent evaluation harnesses to assess response quality, reliability, factual grounding, safety, workflow completion, and business effectiveness.Build agentic workflows and multi-agent systems using LangGraph, LangChain, Claude Code SDK, OpenAI Agent Development Kit (ADK), and related frameworks.Integrate AI agents with enterprise APIs, cloud services, databases, knowledge repositories, vector search, RAG patterns, and internal platforms.Support AI governance, model observability, monitoring, security, responsible AI controls, and operational support for production-grade AI solutions. Must have: 5 to 7 years of overall IT industry experience, with strong hands-on engineering background in cloud infrastructure, platform engineering, DevOps, or related disciplines.3 to 5 years of proven experience in cloud technologies across Azure and AWS.Strong knowledge of Azure and AWS Landing Zone architecture, cloud foundations, account or subscription structures, governance, and security controls.Hands-on experience with Terraform module development, GitHub, GitHub Actions, and CI/CD automation.Practical experience with Kubernetes, containerized workloads, AKS, and EKS.Strong Python programming skills for automation, AI application development, API integrations, and platform engineering use cases.Solid understanding of on-premises to cloud connectivity including VPN, ExpressRoute, Direct Connect, routing, DNS, firewalls, private endpoints, and enterprise network segmentation.Hands-on experience with core Azure and AWS services used to support application, data, platform, and AI workloads.Deep understanding of enterprise networking, identity, access management, observability, and security patterns in public cloud.Strong understanding of LLM fundamentals, generative AI concepts, embeddings, vector search, RAG patterns, model limitations, and responsible AI considerations.Practical experience with prompt engineering, prompt optimization, structured outputs, prompt chaining, and AI workflow design.Hands-on exposure to AI agent development, agent orchestration, AI agent state management, and AI evaluation harnesses.Experience with AI development frameworks and tools such as LangGraph, LangChain, Claude Code SDK, OpenAI ADK, or similar technologies.Ability to collaborate effectively with security, networking, infrastructure, vendor, product, and application development teams in a large enterprise environment. Nice to have: Experience deploying applications and platforms using resilient, highly available, multi-region, and disaster recovery aware architectures.Experience with Azure AI Foundry, Azure OpenAI, AWS Bedrock, Anthropic Claude, OpenAI, enterprise AI gateways, or internal AI enablement platforms.Experience building AI copilots, intelligent assistants, agentic automation workflows, enterprise knowledge assistants, or AI-powered self-service platforms.Familiarity with Model Context Protocol (MCP), AI gateways, vector databases, semantic search, RAG pipelines, and enterprise knowledge integrations.Valid Azure and/or AWS certifications, preferably beyond a single fundamentals exam.Experience working in financial services, regulated environments, or large-scale enterprise technology organizations.