Senior AI Platform Engineer

Teksystems — Hong Kong Sar · Posted ~1 day ago

Senior

Skills

Kubernetes Docker CI/CD Jenkins GitLab CI GitHub Actions Ansible Prometheus Grafana Cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

A senior engineering role focused on operating AI platforms and cloud infrastructure. The position requires expertise in containers, automation, deployment pipelines, monitoring, and production reliability.

Highlights

Work on enterprise AI infrastructure, cloud platforms, automation, and highly available production systems in a technically challenging environment.

Description

Large Financial InstitutionAI PlatformGood Technology Key Responsibilities Support and maintain enterprise AI platforms and cloud infrastructure in production environments.Manage Kubernetes clusters, containerized workloads and platform services to ensure high availability and operational stability.Design, implement and maintain CI/CD pipelines using Jenkins, GitLab CI and/or GitHub Actions.Automate deployment, configuration management and patch management using Ansible Automation Platform (AAP) or Ansible.Perform production releases, change implementation, rollback planning and post-deployment verification.Monitor platform health, application performance and infrastructure using Prometheus, Grafana, CloudWatch and centralized logging solutions.Investigate production incidents, perform root cause analysis (RCA) and implement preventive improvements.Support PostgreSQL databases, container platforms and cloud infrastructure.Collaborate with AI Engineers and Data Engineers to support model deployment, AI services and data platform operations.Maintain operational documentation, standard operating procedures (SOPs) and support records.Drive automation initiatives to improve platform reliability and operational efficiency. Requirements Bachelor’s Degree in Computer Science, Information Technology or related discipline.Minimum 4 years of experience in Platform Engineering, DevOps, SRE, Cloud Infrastructure or Production Support.Strong experience supporting mission-critical production environments.Experience working in Banking, Financial Services, Insurance or other regulated industries is highly preferred.Strong troubleshooting and analytical skills.Good communication skills and ability to work with cross-functional teams.