Staff MaaS Backend Engineer

Bitdeer — Singapore · Posted ~1 day ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

About Bitdeer Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. What You Will Be Responsible For Architecture, Strategy, and Leadership: Co-own the end-to-end MaaS system design with the Principal Architect, authoring decision records and defending technical trade-offs. Drive the platform through its maturity roadmap by delivering operable, measurable capabilities rather than mere demos. Lead technical execution by setting stringent Go and API standards, mentoring engineers, and aligning cross-functional teams.Inference Gateway and API Surface: Own the wire compatibility contract for major formats (OpenAI, Anthropic), supporting advanced eatures like streaming, tool calling, and structured output. Evolve the routing tier to handle load-aware, model-aware, and prefix-cache-aware endpoint selection with robust circuit breaking and fallback mechanisms. Run versioning and deprecation as a published contract to guarantee external customer code stability across underlying changes.Performance Optimization and Model Lifecycle: Maximize platform economics and performance by optimizing token throughput, KV cache tiering, and time-to-first-token (TTFT) latency at the p95/p99 levels. Mature the model serving control plane by integrating deployment tooling, LoRA multiplexing, and cold-start-aware autoscaling directly with the Kubernetes fleet. Treat regressions in cost-per-million-tokens or latency metrics as critical system incidents.Global Topology and Reliability (SLOs): Scale the platform to a globally distributed architecture featuring regional inference pools, capacity-aware failovers, and an active-active control plane. Define, publish, and rigorously defend strict Service Level Objectives (SLOs) baselined against actual system performance rather than aspirations. Ensure operational resilience through peak-concurrency load testing, robust on-call runbooks, and predictable load-shedding during overloads.Security, Identity, and Multi-Tenant Isolation: Enforce fail-closed authorization, robust multi-tenant isolation, and zero-retention data paths across the network, cache, and storage layers. Manage the complete lifecycle of API keys and OAuth credentials while distributedly enforcing rate limits and quotas without relying on client-supplied identifiers. Design strict abuse, rate, and prompt-injection controls, treating all user and model-generated content as untrusted data.Token Metering and Billing CorRectness: Build an idempotent, exactly-once metering system to capture uncached, cached, output, and reasoning tokens accurately across all requests. Maintain a high-volume usage ledger that enforces prepaid spend caps asynchronously and reconciles perfectly with the invoicing system. Safeguard business integrity by treating any metering or billing defect as a critical revenue and trust incident.Safe Migrations and Observability: Execute zero-regression, incremental platform upgrades using strangler-style replacements, shadow traffic, and stateful dual-writes. Deliver end-to-end request tracing and cost telemetry, defining internal schemas for token and quality attributes. Enforce absolute log hygiene by ensuring prompts, completions, and PII are never retained outside of explicit, consented policies. How You Will Stand Out Bring 8+ years of backend engineering experience, including 3+ years owning a high-traffic, multi-tenant API platform for paying customers. Act as a senior technical voice who can collaborate closely with architects, write rigorous design documents, and commit to executing architectural decisions effectively.Demonstrate expertise in scaling distributed systems through multi-region active-active deployments, caching, backpressure, and targeted performance engineering that measurably lowers unit costs. You must have a proven track record of safely executing zero-downtime brownfield migrations for stateful subsystems—like metering or ledgers—without regressions or accounting gaps.Possess deep hands-on proficiency with Go-based services, production Kubernetes (including Envoy and GPU-aware scheduling), and the architectural trade-offs of datastores like PostgreSQL, Redis, and Kafka. Additionally, you will drive operational visibility by owning end-to-end observability strategies using OpenTelemetry and high-cardinality analytics stores.Apply a systems-level understanding of LLM serving to manage complexities like server-sent-event streaming, KV/prefix caching, and the trade-offs between time-to-first-token (TTFT) and throughput. Leverage this foundation to build highly reliable, exactly-once metering and billing systems that accurately reconcile billions of events under partial failure conditions.Enforce strict multi-tenant security disciplines by designing fail-closed authorization, mandating verified identities, and guaranteeing absolute cross-tenant isolation. Bring operational maturity to a revenue-bearing platform by carrying on-call responsibilities, running blameless incident reviews, and translating outages into structural improvements. What You Will Experience Working With Us A culture that values authenticity and diversity of thoughts and backgrounds;An inclusive and respectable environment with open workspaces and exciting start-up spirit;Fast-growing company with the chance to network with industrial pioneers and enthusiasts;Ability to contribute directly and make an impact on the future of the digital asset industry;Involvement in new projects, developing processes/systems;Personal accountability, autonomy, fast growth, and learning opportunities;Attractive welfare benefits and developmental opportunities such as training and mentoring. Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.