Principal Engineer - Platform and Reliability

Skillflotalent — Indonesia (for Vietnam branch/hub) · Posted ~2 hours ago

Lead Full-time

Skills

Software architecture Platform engineering Reliability engineering Cloud infrastructure Distributed systems Cloud Distributed Systems Platform Engineering

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A principal engineer is sought to lead platform and reliability initiatives, make architectural decisions, and build scalable technology solutions in a fast-moving engineering environment.

Highlights

Principal-level engineering opportunity with strong ownership, architectural influence, and work on ambitious technology products. The role emphasizes technical depth and engineering impact.

Description

Principal Engineer, Platform & Reliability About SkillfloTalent SkillfloTalent is a talent solutions firm based in Ho Chi Minh City, Vietnam, specializing in connecting exceptional candidates with high-growth technology companies across Southeast Asia. We combine deep industry knowledge in AI, Cloud, and enterprise technology with a hands-on, relationship-driven approach to recruitment. Our mission is to match the right people with the right opportunities — enabling both candidates and clients to grow faster. Company Information This is a fast-moving, early-stage technology company building ambitious products at the intersection of real-time communication, AI, and enterprise software. The engineering culture prizes pragmatism, ownership, and sound judgment as much as technical depth. You will join a small, high-caliber team where your architectural decisions will have immediate and lasting impact on the product and its customers. The company operates at globally competitive engineering and operational standards. Job Summary We are seeking a Principal Engineer to own the platform and reliability function at a critical stage of growth. You will make foundational architectural decisions, lead incident response, and raise the bar for how the team builds and operates customer-facing production systems. This is a hands-on leadership role suited to someone who thrives under the constraints and opportunities of an early-stage environment. Key Responsibilities Design and evolve the architecture of a secure, scalable, multi-tenant SaaS platform through a significant growth phase.Own production reliability end-to-end, including incident response, root-cause analysis, and subsequent system improvements.Make sound architectural decisions under constraints of time, cost, and incomplete information.Define and enforce strong authorization and isolation boundaries across the multi-tenant platform.Lead the design and implementation of distributed systems components including queues, concurrency, state management, retries, idempotency, backpressure, consistency, and failure recovery.Manage cloud infrastructure, serverless systems, containers, CI/CD pipelines, secrets management, and production observability tooling.Coordinate zero-downtime migrations and releases across multiple services.Contribute hands-on software engineering using TypeScript, Node.js, and modern backend systems.Collaborate closely with product and engineering stakeholders, communicating architectural tradeoffs clearly in written and verbal form. Requirements A strong record of building and operating customer-facing production systems.Direct experience scaling a multi-tenant SaaS, infrastructure, realtime, or AI-enabled product through a significant growth phase.Experience working in a high-growth North American or globally operated technology company, or in an environment with comparable engineering and operational standards.Deep knowledge of distributed systems, including queues, concurrency, state management, retries, idempotency, backpressure, consistency, and failure recovery.Experience designing secure multi-tenant platforms with strong authorization and isolation boundaries.Experience with cloud infrastructure, serverless systems, containers, CI/CD, secrets management, and production observability.Strong hands-on software-engineering ability, ideally with TypeScript, Node.js, and modern backend systems.Personal ownership of production incidents, root-cause analysis, and subsequent system improvements.Demonstrated ability to make sound architectural decisions under constraints of time, cost, and incomplete information.Strong written and verbal communication skills.Currently operating at Staff, Principal, Founding Engineer, or senior Technical Lead level in a high-growth technology company.