AIOps Engineer - AI Infrastructure and Orchestration

T Mobile Polska — Poland · Posted ~5 hours ago

Senior Full-time

Skills

vLLM OpenShift Kubernetes GPU Infrastructure NVIDIA GPU Management Model Lifecycle Management Container Orchestration Horizontal Pod Autoscaling Infrastructure Automation NVIDIA GPUs Hugging Face Enterprise Amazon S3 HPA

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology-driven organization is seeking an AIOps Engineer specializing in AI infrastructure and orchestration. You will design, deploy, and maintain high-performance model inference services on Kubernetes-based platforms running on bare-metal GPU infrastructure. Responsibilities include managing NVIDIA GPU resources, automating model lifecycle operations, implementing workload-aware autoscaling, and optimizing AI serving environments for multiple models and tenants.

Highlights

Work at the forefront of AI infrastructure, GPU computing, and cloud-native orchestration. Gain exposure to advanced inference systems, automated model deployment, resource optimization, and emerging technologies shaping next-generation connectivity.

Description

Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation. You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity. Responsibilities WHAT TASKS AWAIT YOU? Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.Maintain audit logging capabilities to support compliance, security investigations, and operational governance.Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives. WHAT SKILLS WILL BE APPRECIATED? 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.Strong experience with Kubernetes and OpenShift administration in production environments.Proven experience deploying and operating vLLM-based inference platforms in production.Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.Strong Python programming skills with experience developing automation and operational tooling.Experience with Bash scripting and Linux systems administration.Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.