Description
About FirstPrinciples
FirstPrinciples is a research company building AI for scientific discovery.
It began with Theo, the AI Physicist, and has since grown into a product-focused company with two additional systems including Theo Conjecture, built for automated conjecturing, and Theo Collaborator, an adaptive environment for doing complex research with AI.
We're a fast-growing, remote-first team of builders, researchers, engineers, and thinkers working across Canada, the US, the UK, and expanding globally.
What brings us together is a shared curiosity about how the universe works, and a belief that we can build systems that help us explore it more effectively.
We spend our time working on questions that don't have clear answers, like how to design AI that can reason through scientific problems, and how the scientific process as a whole might evolve.
This is work that sits somewhere between creativity and rigorous thinking, and often requires comfort with ambiguity and iteration.
If you're someone who enjoys tackling big, abstract problems and exploring ideas that don't yet have a defined path forward, you'll likely find the work here interesting.
Why This Role Exists
We’re building the next generation of infrastructure for AI-driven scientific discovery, and we need someone who can help own the systems that make our research and inference workloads reliable, scalable, and fast.
This role is about building and operating the compute foundation behind our AI Physicist: Kubernetes clusters, Linux systems, GPU infrastructure, cloud environments, HPC-style compute, deployment workflows, monitoring, and automation.
As our workloads grow, we need infrastructure that can support both experimentation and production-like inference across cloud, bare metal, and hybrid environments.
You’ll play a central role in shaping how we run compute at FirstPrinciples.
That includes provisioning and managing clusters, improving reliability and observability, reducing operational toil, supporting researchers and engineers, and helping us make practical decisions about when to use managed cloud services, self-managed Kubernetes, Slurm-style systems, or owned hardware.
We’re looking for someone hands-on, systems-oriented, and comfortable working in a fast-moving research environment.
You should have strong Kubernetes and Linux fundamentals, good operational instincts, and enough experience with cloud and HPC/GPU infrastructure to help us build toward a robust bare metal and multi-cloud inference platform.
What You’ll Do
Design, deploy, and operate Kubernetes infrastructure for AI inference, research, and engineering workloadsSet up and manage GPU and HPC-style compute environments, including scheduling, utilization, job management, and node-level troubleshootingWork with systems such as Kubernetes, Slurm or similar schedulers, container runtimes, GPU drivers & libraries (ie; CUDA), storage systems, and observability toolsBuild and manage Linux-based compute environments, including provisioning, networking, storage, monitoring, access control, and lifecycle managementHelp architect bare metal, cloud, and hybrid infrastructure across AWS, GCP, Azure, or equivalent platformsOwn the reliability and operational health of infrastructure systems, including monitoring, alerting, incident response, capacity planning, and performance tuningImprove deployment workflows, automation, configuration management, secrets management, and infrastructure-as-code practicesPartner with ML engineers, researchers, and software engineers to understand workload requirements and translate them into practical infrastructure designsEvaluate tradeoffs between managed cloud services, self-managed Kubernetes, HPC schedulers, bare metal deployments, and multi-cloud architecturesBuild tooling, documentation, runbooks, and operational practices that help the team move quickly without making infrastructure fragile or opaqueBalance speed and robustness, knowing when to prototype quickly and when to harden systems for long-term use
Who You Are
Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systemsDeeply comfortable with Linux administration, including debugging networking, storage, system services, permissions, performance issues, and node-level failuresExperienced with Kubernetes in real environments, including cluster operations, deployments, networking, observability, scaling, and troubleshootingComfortable working with cloud infrastructure on AWS, GCP, Azure, or equivalent platformsFamiliar with infrastructure automation and configuration tools such as Terraform, Ansible, Helm, ArgoCD, GitOps workflows, or similar systemsExperienced with GPU-heavy, compute-heavy, or HPC-style workloads, especially in environments involving AI, ML, research computing, or scientific workloadsAble to work across bare metal and cloud environments, and interested in the practical tradeoffs between the twoComfortable reasoning about resource scheduling, cluster utilization, autoscaling, storage, networking, and observability for distributed workloadsPractical and ownership-oriented; you can take ambiguous infrastructure needs and turn them into working systemsComfortable collaborating across disciplines, especially with researchers and engineers who may not think in infrastructure termsAble to operate independently as a senior or strong intermediate contributor, while knowing when to bring others into important technical decisionsMotivated by building foundational systems that make ambitious technical and scientific work possible
Bonus
Hands-on experience with production-grade LLM inference and serving engines, such as vLLM, SGLang, or TensorRTExperience working at an AI company, ML infrastructure team, research lab, university compute environment, HPC center, or scientific computing organizationExperience supporting model inference, model serving, distributed training, high-throughput batch workloads, or internal ML platformsHands-on experience with Slurm or similar HPC schedulers, including job scheduling, resource allocation, queue management, and cluster configurationExperience operating GPU infrastructure, including NVIDIA drivers, CUDA, container runtimes, scheduling, utilization, and hardware failure modesExperience with RDMA, InfiniBand, high-performance networking, distributed filesystems (ie.
Lustre, BeeGFS), object storage, or storage systems for compute-heavy workloadsExperience with Kubernetes operators, custom controllers, CRDs, or platform tooling for AI/ML workloadsExperience with Prometheus, Grafana, Loki, OpenTelemetry, Datadog, or similar monitoring, logging, and observability toolsExperience with container registries, image optimization, CI/CD systems, deployment pipelines, and secure software deliveryExperience leading engineering operations or infrastructure efforts while remaining hands-on technicallyFamiliarity with security, access control, secrets management, and reliability practices in production or research environments
What You’ll Get
The opportunity to work on foundational problems at the intersection of AI and physicsA high-trust, low-bureaucracy environment with real ownershipRemote-first work with flexibility in how you structure your dayExposure to cutting-edge ideas across AI, scientific discovery, infrastructure, and emerging technologiesA culture that values curiosity, depth of thinking, and first-principles reasoningThe chance to shape the compute and inference infrastructure behind advanced AI systems for scientific discovery