Senior Backend Engineer

Gamania Digital Entertainment — Taiwan · Posted ~22 hours ago

Senior

Skills

backend engineering system architecture low-latency services high-availability systems real-time streaming relational databases NoSQL databases Kubernetes LLM optimization GPU deployment LLMs NoSQL vector retrieval GPU streaming services

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior backend engineering role focused on building highly available, low-latency infrastructure for AI-driven services. You will shape system architecture, develop real-time streaming and event protocols, optimize relational and NoSQL data retrieval, reduce LLM costs, troubleshoot Kubernetes at scale, and deploy models on GPU infrastructure.

Highlights

Work on scalable AI-agent infrastructure, design low-latency real-time services, build long-term AI memory systems, optimize LLM costs, and operate GPU-backed workloads at high concurrency.

Description

We are seeking a Senior Backend Engineer with passion to bring our products to success. The ideal candidate will support the AI agent platform backend team by reviewing system architecture design, and spearheading problem-solving initiatives. Design and maintain low-latency, high-availability streaming services and event protocols that support real-time integration between client applications and AI agentic services at scale.Architect and optimize state management and retrieval across relational and NoSQL document stores to support long-term AI memory, including vector-based semantic retrieval.Reduce LLM token cost through prompt restructuring, caching of repeated context, and consolidating the number of model calls per conversation turn.Diagnose and resolve Kubernetes resource bottlenecks under high concurrency, covering resource allocation, health checks, and node pool planning.Deploy distilled models on GPU node pools, and manage model routing, load balancing, rate limiting, and failover through an LLM gateway. Bachelor's degree in Computer Science or equivalent practical experience.2+ years of experience in backend infrastructure, API design, and distributed systems development.Experience with Python, including asynchronous programming, streaming response handling, connection pooling, and debugging event loop behavior in production.Experience operating relational databases, distributed caching layers, and containerized services in production.Experience operating Kubernetes in production.Experience building, deploying, and maintaining high-concurrency services on GCP or AWS. Experience running managed Kubernetes, managed relational databases, message queues, container registries, and centralized logging in production, with a track record of reducing cloud infrastructure cost.Experience operating LLM applications in production, including latency and cost governance, caching strategies, streaming responses, and tool/function calling.Experience implementing retrieval-augmented generation systems or vector databases at scale, with an understanding of vector similarity search trade-offs.Experience designing or implementing agentic workflows, multi-agent systems, or the Model Context Protocol.Experience serving self-hosted models on GPU infrastructure, including scheduling, and deploying quantized or distilled models with high-throughput inference engines.