Production Support Engineer - LLM/AI Platform

Tribus Technology — Hong Kong Sar · Posted ~4 hours ago

Senior Full-time Onsite Visa History ✓

Skills

Production support Troubleshooting AWS Kubernetes Cloud infrastructure Distributed systems SRE AI infrastructure LLM platforms Monitoring Automation Scripting LLM Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An onsite Production Support Engineer is needed to operate and improve an enterprise LLM and AI tooling platform. You will troubleshoot production incidents, investigate latency, availability, API, capacity, and regional performance issues, support cloud and Kubernetes infrastructure, and build automation that improves diagnostics. The role sits between production engineering, SRE, and AI infrastructure.

Highlights

Hands-on production engineering role at the intersection of SRE, AI infrastructure, and cloud systems. It offers exposure to LLM technologies, distributed systems, Kubernetes, monitoring, automation, and performance optimization while working on infrastructure used broadly by technical teams.

Description

Production Support Engineer - LLM / AI Platform Hong Kong | Onsite We’re working with a leading global quantitative investment manager where technology, data and engineering are fundamental to how the business operates. As investment in internal AI capabilities continues to grow, they’re looking for a Production Support Engineer to help operate and improve the platform that provides engineers and users across the firm with access to LLMs and AI tooling. This is a hands-on role sitting between Production Engineering, SRE and AI infrastructure, with exposure to modern cloud, distributed systems and LLM technologies. What you’ll be doing Own and troubleshoot production issues across the LLM platform and model-provider integrationsInvestigate latency, availability, API, capacity and regional performance issuesSupport AWS/Kubernetes infrastructure and developer toolingBuild scripting and automation to improve diagnostics and reduce manual workImprove monitoring, dashboards and alertingSupport releases and production changesWork with engineering, infrastructure teams and external providers to resolve incidents What we’re looking for Production Support, SRE or Platform Operations experience with genuine ownership of production incidentsExperience supporting LLM/AI infrastructure OR within trading/hedge fundsStrong Linux and Windows knowledgePython, Bash and/or PowerShell scriptingAWS, Kubernetes and distributed/API-driven systemsMonitoring and observability —Grafana, Prometheus or similarSQL experience with PostgreSQL, SQL Server or similarStrong communication and incident-management skills Useful additional experience LLM gateways, inference APIs, AWS Bedrock, streaming or prompt cachingAI coding assistants or developer toolingTerraform / AnsibleMulti-region production environments A strong opportunity for an experienced Production Engineer/SRE to work closer to AI and LLM infrastructure within a highly technical quantitative investment environment.