Senior Cloud Platform Engineer

Xora Innovation โ€” United States ยท Posted ~1 day ago

Senior Hybrid

Skills

Cloud infrastructure Kubernetes Infrastructure as Code Terraform AWS GCP Azure Backend services APIs CI/CD Release engineering Observability Prometheus Grafana OpenTelemetry Go Python Rust Service-level objectives Reliability engineering Containerization Security Helm ArgoCD Flux Slurm GitOps

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Build the infrastructure foundation behind a sophisticated AI and computational platform as a senior cloud platform engineer. You'll own Kubernetes-based infrastructure, backend services, deployment automation, observability, reliability, security, and reproducible installations across diverse customer environments. This is a hands-on opportunity for an engineer who enjoys solving complex platform problems in an early-stage setting.

Highlights

Highly hands-on senior role with broad ownership of cloud infrastructure, platform services, deployment, observability, and reliability. Offers the chance to build foundational systems in an early-stage environment, work across cloud and advanced computing workloads, and make meaningful architecture and engineering decisions.

Description

About Elemynt ELEMYNT is an early-stage startup built by Xora Innovation. We develop applied intelligence that brings AI into the real world. Our platform combines advanced machine learning, high-performance simulation, and modern software engineering to accelerate the design, validation, and deployment of new materials. Our work sits at the intersection of AI, physics, and large-scale computation. The problems are hard, the stakes are high, and the impact is tangible. About The Role This role builds what the rest of Elemynt's engineering runs on: the cloud infrastructure, the core platform services, and the observability that keeps the product reliable. The platform runs wherever each customer chooses, whether that's their cloud, their own compute, or a hybrid, and usually inside environments they operate themselves. That's the hard part. The foundations you build have to be reproducible, measurable, and carry their own security wherever they land. This is a deeply hands-on, build-focused senior role. In this role, you'll design and own the cloud substrate, the services at the core of the product, and the metrics, logs, and traces that keep all of it visible. Those building blocks are what gets packaged and deployed wherever the platform needs to run, so the care you put in shows up in every install. How fast and how safely the whole company can ship depends on how solid these foundations are. What You Will Do Build and own the cloud infrastructure foundations (networking, identity and access, Kubernetes, and infrastructure-as-code) that every service and workload runs on. Design and build the core platform services and internal APIs the product is made of, meant to run reliably wherever it's deployed. Own the cloud deployment process end to end: the pipeline that turns platform services into versioned, reproducible artifacts and rolls them out with staged releases and clean rollbacks. Stand up the observability layer: metrics, logs, traces, dashboards, and alerting that make a failing service quick to find and diagnose. Instrument service-level objectives and health signals so reliability is measurable and regressions show up before they reach a customer. Keep the foundations portable and reproducible, so the platform stands up the same way across every environment it runs in. Produce the deployment-ready building blocks (container images, Helm charts, infrastructure-as-code modules) that make installs and upgrades clean and repeatable wherever the platform runs. Harden the platform's foundations: secrets, certificate handling, and network boundaries that protect the software and its data wherever it runs. What We Are Looking For Bachelor's or Master's degree in Computer Science or a related engineering field, and 6+ years building and shipping production software, with real depth in cloud infrastructure and platform engineering. Deep hands-on experience building cloud infrastructure with Kubernetes and infrastructure-as-code (Terraform or similar) on at least one major cloud (AWS, GCP, or Azure). Experience designing and operating core backend services and APIs that other engineers and systems depend on. Direct ownership of a cloud deployment process: build and release pipelines, versioned artifacts, and safe, reproducible rollouts. Hands-on experience building observability into production systems (metrics, logs, and traces) and using it to debug real incidents (Prometheus, Grafana, OpenTelemetry, or similar). Strong software-engineering fundamentals and hands-on coding in a systems or backend language (Go, Python, Rust, or similar): this role builds the platform in code. Experience defining service-level objectives and designing for reliability, making systems observable and reproducible from the start. Comfort building foundational systems others depend on in an early-stage, ambiguous environment, making sensible scope, speed, and quality trade-offs. NICE TO HAVE Experience building platform components that run across varied deployment environments, including customer-controlled ones. Experience running workloads across more than one runtime: cloud Kubernetes plus HPC schedulers (Slurm or similar) or bare metal. GPU scheduling or multi-tenant cluster experience. GitOps and progressive-delivery patterns (ArgoCD, Flux, staged rollouts). Experience packaging or serving ML models, or supporting ML and data workloads on a shared platform. Exposure to scientific computing, simulation, or other large-scale technical workloads. LOCATION Singapore or United States. We're hiring in both to reach the right person. Work model is on-site or hybrid, set per location. CLOSING NOTE If you don't tick every box but this is clearly your kind of work, get in touch.