AI DevOps Specialist

Celestica — Canada · Posted ~3 hours ago

Senior Full-time Hybrid

Skills

DevOps cloud platforms AI platforms secure infrastructure CI/CD Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AI-focused DevOps position responsible for deploying, securing, and optimizing development platforms. The role bridges AI workflows, cloud services, secure networking, and engineering toolchains.

Highlights

Hybrid role combining AI enablement, cloud operations, security, and modern engineering practices.

Description

Job Title: AI DevOps Specialist (AI Enablement & Secure DevOps) Remote Position: Hybrid Region: Americas Country: Canada State/Province: Ontario City: Toronto Project Objectives & Role Summary The AI DevOps Specialist is primarily responsible for the technical deployment, secure enablement, administration, and continuous optimization of Celestica’s global HPS Artificial Intelligence (AI) sandbox, modeling, and software engineering toolchains. This specialist will bridge the gap between AI development pipelines, secure networking infrastructure, and cloud platform services. They will take a hands-on lead in configuring foundational AI services within the DevOps space. implementing agentic workflow architectures, ensuring secure, uninterrupted developer access to AI coding assistants (such as Claude Code, Gemini Antigravity Code, and Codex), and establishing robust cloud governance. Core Responsibilities & Scope of Work 1. AI Sandbox Platform Administration & Deployment Multi-Phase Sandbox Rollout: Own the deployment and lifecycle management of the HPS AI Sandbox environment across cloud and hybrid infrastructures.Agentic Frameworks: Establish and configure platforms such as Claude Code, Gemini Antigravity Code, and Codex for the design and orchestration of agentic AI workflows and LLM-backed applications.Project Governance: Oversee the HPS-designated AI projects, including security permissions, service accounts, and IAM roles.Model Enablement & Tuning: Coordinate the provisioning and scale-out of advanced foundational models. Actively manage, troubleshoot, and resolve API restrictions with cloud providers. 3. Secure DevOps & Network Integration (Zscaler & Firewall) Lab Network Troubleshooting: Diagnose and resolve intermittent connection issues between HPS Design Labs (such as the Innovation Lab) and external AI resources or code repositories (e.g., troubleshooting DNS, routing timeouts, and SSL inspection blocks on GitHub).Traffic Rules & Proxy Controls: Partner with Network Security to define, test, and troubleshoot Zscaler ZTNA app connectors, firewall rules, and proxy exceptions necessary to enable secure outbound AI API traffic while protecting proprietary codebase egress.CI/CD Integration: Work alongside DevOps administrators to embed automated vulnerability checks, binary scanning, and AI-assisted testing steps in Azure DevOps, Jenkins, and GitHub pipelines. 4. Observiability and Monitoring Controls with AI Monitoring: Design and implement rigid monitoring, consumption alerts, and attribution controls in all on-prem and virtualize all environments to track HPS developer usage.Live Auditing: Develop real-time observability using AI dashboards to track consumption, model call costs, and sandbox compute runtimes. 5. AI Security, IP Protection & Compliance Data Sovereignty Compliance: Enforce enterprise policies ensuring that no proprietary hardware schematics, PCB layouts, firmware source code, or IP are ingested into public training models.Evaluation & Testing: Support the isolation of the HPS Innovation Lab to evaluate new open-source models, libraries, and AI security evaluation tools prior to general HPS rollout.Education & Experience Bachelor’s degree in Computer Science, Software Engineering, DevOps, Cloud Engineering, or equivalent technical experience.4+ years of hands-on experience in DevOps, Cloud Engineering, or System Administration, with at least 2 years focused specifically on AI/MLOps platform delivery.Knowledge/Skills/Competencies Required Technical Skills AI Toolchain & LLM Tooling: Proven experience deploying and maintaining containerized mircoservices, and local/cloud LLM APIs.Networking & Security Engineering: Solid understanding of enterprise networking protocols (DNS, TCP/IP routing, NAT, SSL/TLS handshake) and secure access controls (Zscaler ZTNA, enterprise Firewalls).Containerization & Orchestration: Strong proficiency with Docker, docker-compose, and Kubernetes to deploy scalable sandbox services.Code Assist Integration: Familiarity with the configuration of developer-focused AI integrations like Claude Code, Gemini Antigravity inside Linux CLI environments.Automation & Scripting: Strong scripting abilities in Python (specifically utilizing AI/ML libraries, request handling, and GCP SDKs) and Bash. Preferred Certifications Linux Enterprise Red Hat (RHCE / RHCSA)GCP/AWS EngineeringCertified Kubernetes Administrator (CKA)HashiCorp Certified: Terraform Associate Working Style & Competencies Diagnostic Mindset: Exceptionally strong debugging skills for networking, package distribution, and cloud service interconnections.Proactive Collaboration: Able to work cross-functionally with HPS Hardware Design Teams, Enterprise IT, Corporate Security, and external consulting suppliers (e.g., Elastify).Detail-Oriented Documentation: High commitment to writing complete, clear Standard Operating Procedures (SOPs), system topologies, and budget management guidelines.Physical Demands Salary The stated range includes Base Salary and target Short-Term Incentive (STI) compensation only. A comprehensive benefits package is offered in addition to this range. The range described in this posting is an estimate by the Company, and may change based on several factors, including but not limited to a change in the duties covered by the job posting, or the credentials, experience or geographic jurisdiction of the successful candidate. 109,000 - CAD 173,000