Description
About the roleWasteFlow runs computer-vision sensors on waste-processing lines โ a fleet of on-prem, PLC-connected edge devices that stream detections and drive automation (conveyor control, alarms, sorting).
Behind them sits a TypeScript product, a Postgres/TimescaleDB system of record, and a full observability and infrastructure-as-code platform on AWS.
We're hiring a Senior Platform / DevOps Engineer to own that platform end to end.
This is a role with a single, clear owner: you will run our cloud infrastructure, our edge-fleet deployment system, our CI/CD, our databases, and our observability stack โ and set the standards the rest of the team builds against.
You'll be the person we trust to keep production stable while making it dramatically easier to ship.
The center of gravity is platform and DevOps.
You'll also own the plumbing of our MLOps โ the MLflow model registry, artifact tracking, and how trained models get provisioned and pushed to sensors in the field โ working closely with the ML team, who own the model logic itself.
What you'll ownCloud infrastructure & IaC.
Our AWS footprint managed as code with OpenTofu; GitHub repositories and org scaffolding provisioned with Pulumi; sensor and host configuration automated with Ansible.
You'll keep infrastructure reproducible, versioned, and reviewable โ no snowflakes.
Edge fleet deployment.
Our "MLHub / Control Center" pipeline โ CLI and API driving an Ansible-orchestrated flow (Configure โ Install โ Provision โ Deploy โ Integrate โ Initialize) out to devices in the field, pulling from GitHub, the model registry, and config stores.
Edge software ships as versioned Python packages (built and published by CI, installed and updated on devices by Ansible) rather than containers.
You'll harden this into reliable, observable, one-command provisioning and over-the-air updates across a growing fleet, connected over a Tailscale mesh.
CI/CD & developer platform.
Our multi-repo setup (a core repo plus standardized sub-repos) and the automated workflows that guard it: format/lint, static analysis, automated tagging, test, and package publishing โ versioned Python packages to our internal registry are the primary release artifact for edge software.
You'll own repo standardization (pyproject.toml, packaging and versioning config, Dockerfiles for cloud services, editor and CI config), keep pipelines fast and green, and make "create a new service" a paved road.
Cloud compute & containers.
Our cloud services run as Docker containers on AWS ECS / Fargate โ build, ship, run, scale, and roll back predictably.
(Edge devices are not containerized; they run versioned Python packages, as above.)
Databases & storage.
PostgreSQL 17 with the TimescaleDB extension for sensor and detection time-series, on AWS RDS, plus S3 object storage.
You'll own reliability, backups, migrations, access, and performance as data volume grows.
Observability & reliability.
Our Grafana Cloud stack: Prometheus metrics, Loki logs, Grafana Alloy and Fluentd collection, Grafana Faro for frontend, CloudWatch, and K6 load testing โ with service discovery for a dynamic fleet and alerting into Slack.
You'll drive SLOs, on-call hygiene, and getting to root cause fast when a sensor in a facility misbehaves.
Security.
Secrets and credentials in HashiCorp Vault, network access over Tailscale, and sensible least-privilege practices across cloud and edge.
MLOps plumbing (in partnership with the ML team).
MLflow model registry and artifact tracking (models, configs, runs), and the provisioning/OTA path that gets a validated model onto the right sensors.
You'll build and operate this infrastructure; the ML team owns model pipelines and the quality/automation gate.
Together you'll push our deploy-time down from ~15 days toward same-day, plug-and-play deployment.
What success looks like (first 6โ12 months)Model/artifact tracking (MLflow) productionized and adopted as the single source of truth for models, configs, and runs.Fleet provisioning and OTA turned into a repeatable, observable, low-touch process as sensor count scales.Time-to-deploy a model to a sensor cut materially (our roadmap targets 5 days/sensor by end of 2026, trending toward under a day) โ the infrastructure half of that is yours.No single point of failure on infra, DB, or deployment: everything documented, reproducible, and runnable by more than one person.Must-have2โ5 years of professional experience, including at least one project where you set up and operated a live, running production infrastructure end to end โ plus hands-on exposure to building an ML model and/or active software development.Strong, hands-on DevOps / platform engineering background โ you've owned production cloud infrastructure (ideally AWS) as the person accountable for it.Infrastructure as code in anger โ Terraform/OpenTofu, Pulumi, and/or Ansible โ plus solid CI/CD pipeline design (GitHub Actions or equivalent).Python packaging & distribution โ building, versioning, and publishing installable Python packages (setuptools/pyproject.toml, internal package registries) as a deployment mechanism.Containers for cloud services: Docker plus an orchestration/compute platform (ECS/Fargate, Kubernetes, or similar).Observability with the Grafana/Prometheus ecosystem (or close equivalents) โ metrics, logs, dashboards, alerting.Comfortable operating relational databases in production (PostgreSQL preferred), including backups, migrations, and performance.Linux fluency and comfortable scripting in Python and/or Bash.Able to work independently and set the standard โ this role owns the platform, it doesn't wait to be told how.Nice-to-haveEdge / IoT / fleet experience: managing many remote devices, OTA updates, mesh networking (Tailscale/WireGuard).Industrial / hardware exposure: PLCs (Siemens S7 / Snap7), MQTT (Mosquitto), device I/O, real-time constraints.MLOps tooling: MLflow, Weights & Biases, model registries, model serving/provisioning.Familiarity with our product stack (TypeScript, Fastify, Next.js, Supabase) โ enough to collaborate, not to own the frontend.TimescaleDB or other time-series databases; gRPC.