AI Engineer

Black Nova Vc — Australia · Posted ~2 hours ago

Skills

Software engineering Go Python Distributed systems LLM-based agents AI agent development Evaluation frameworks Structured output Checkpoint and resume systems Production systems LLM

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An AI engineer is sought to build and harden a production-grade agent runtime that can analyze application behavior, validate findings, and prepare fixes. You will design reasoning agents, develop structured execution and budget controls, support long-running resumable workflows, and build evaluation systems with regression tracking. Strong Go or Python skills and experience operating LLM-based agents in production are expected.

Highlights

Work on challenging AI agent and distributed-systems problems, owning agents end-to-end from detection through validation and fix generation. The role emphasizes production engineering, measurable evaluation, reliability, and meaningful autonomy.

Description

We're building the agent runtime that finds, validates, and prepares fixes for vulnerabilities automatically, while people retain judgment and merge control. That means solving real distributed-systems and reasoning problems, not gluing together prompts. What You'll Do Build and harden Nullify's multi-provider agent runtime — turn loops, structured output, budget governance, checkpoint/resume for multi-hour runs.Design agents that reason about business logic and real application behavior, not just pattern-match against known signatures.Build and grow our evaluation harness — labelled test cases, scoring, regression tracking — so we can measure whether an agent got better or just got different.Own agents end-to-end: detection, validation, fix generation, and the judgment calls in between. What You Bring Strong software engineering fundamentals — Go or Python, comfortable in production systems, not just notebooks.Experience building or operating LLM-based agents beyond a single prompt-response call: tool use, multi-step reasoning, retries, evals.A scientist's instinct for measuring whether something actually works before shipping it.Genuine interest in security — you don't need to be an expert yet, but you need to want to become one.