Summary
✨ AI‑Generated
Join an engineering team improving the reliability and performance of AI-powered voice interactions. You will work across backend services, real-time voice agents, infrastructure, observability, and evaluation systems, establishing measurable quality standards and turning quantitative and qualitative findings into engineering improvements.
Highlights
Work across the full lifecycle of AI voice systems, build measurable evaluation frameworks, improve reliability and latency, and directly shape the quality of real-time conversational experiences.
Description
We are looking for a Full-Stack Software Engineer to improve the quality, reliability, and performance of Frank’s AI-powered voice interviews.
You will work across backend services, real-time voice agents, infrastructure, observability, and evaluation systems.
Your primary responsibility will be to understand how an AI voice interview behaves end to end, establish measurable quality standards, evaluate interviews using quantitative and qualitative metrics, and turn findings into engineering improvements.
What you will do
Develop a deep understanding of the complete voice-interview lifecycle, from participant admission and bot greeting through conversation, transcription, recording, and analysisDesign repeatable evaluation frameworks for comparing prompts, models, voice pipelines, providers, and infrastructure changesDefine, collect, and analyze quantitative metrics, including:Response latency and time to first audioEnd-of-utterance and STT finalization latencyInterruption and false-interruption ratesTranscript accuracy and missing or duplicated turnsInterview completion, failure, and abandonment ratesProvider and pipeline error ratesAudio quality, packet loss, jitter, and connection stabilityToken usage, provider cost, CPU, and memory consumption Evaluate qualitative interview characteristics, including:Question relevance and research-goal coverageQuality and depth of follow-up questionsContext retention and conversational coherenceNeutrality and avoidance of leading questionsNatural turn-taking, pacing, tone, and voice qualityEmpathy and responsiveness to participantsHallucinations, repetition, awkward transitions, and premature endings Build automated evaluation tools using transcripts, telemetry, recordings, LLM-as-judge
Tech you will use here
AI/LLM: OpenAI/Anthropic/others; embedding stores (OpenSearch/pgvector/Pinecone); vectorization & reranking; eval/guardrailsRuntime: TypeScript/Node or Python (both welcome)Tooling: Terraform/CDK (nice-to-have)
You might be a fit if you have
Strong software engineering experience with Python and/or TypeScript.
Experience building, testing, or evaluating LLM-powered applications.
Understanding of real-time voice systems, including STT, LLM, TTS, voice activity detection, endpointing, and interruption handling.Strong analytical skills and experience turning product-quality questions into measurable metrics.Experience with experimentation, statistical analysis, and structured qualitative review.Ability to diagnose problems across application code, external AI providers, media infrastructure, and observability data.Experience writing automated tests and maintaining production-quality systems.Strong written communication and documentation skills.
This is a Senior / Lead-level role reporting directly to the CTO, with a lot of ownership over how we build and ship AI features.
Nice to have
Experience with LiveKit, WebRTC, or similar real-time media technologies.
Experience with Langfuse and OpenTelemetry.
Familiarity with Gemini, OpenAI Realtime or OpenRouter.Experience designing LLM-as-judge evaluations and human-annotation rubrics.Knowledge of conversational research, user interviewing, or qualitative research methods.Experience with AWS services such as ECS, CloudWatch, S3, and SQS.Familiarity with NestJS, PostgreSQL, and infrastructure as code.
DBT/analytics chops; safety/compliance (SOC-2, GDPR DPIAs).
Why you will love it
Own meaningful AI features end-to-end.Pragmatic stack with room to choose the right tool (framework-light where it helps).Fast decisions, real users, real impact.