Summary
β¨ AIβGenerated
Lead the architecture and development of a production-grade conversational AI backend spanning edge and cloud environments. You will design orchestration services for LLMs, tools and agents, while helping evolve streaming speech pipelines, multi-intent handling, safety controls and reliable operations. Strong experience with Python, Azure and edge or on-device AI is expected.
Highlights
Work on a production-scale conversational AI platform with modern cloud, LLM, agent, and edge technologies. The role offers technical leadership, continuous feature development, and exposure to a large deployed user base.
Description
Project description
Our client is advancing its in-vehicle voice assistant into an intelligent, AI-powered companion.
Large-language-model capabilities (Azure OpenAI / ChatGPT) have been running in production across vehicles.
The goal of the project is to develop a backend which is the cloud AI orchestration service behind this: it receives requests from the vehicle, routes them, orchestrates the LLM, tool services and agents, and returns an answer or action to the car.
DXC Luxoft serves as the end-to-end delivery partner, working in a joint product team with the client's engineers on the Azure platform.
This is the series development and operations work package - a live platform serving a large vehicle fleet, which extends sprint by sprint the backend features while availability and backward compatibility are maintained.
New capability in the pipeline includes streaming across the full ASR β LLM β TTS chain, barge-in, multi-intent handling, a guardrails layer for deterministic vehicle-safe answers, agent routing and new tool integrations.
The role is Technical Lead and Location Lead for the engineering team.
It is a hands-on delivery leadership role: the person is accountable for what the team ships, for the service running inside its availability and incident targets, and for the technical growth of the location.
Responsibilities
Own the system architecture for hybrid (cloud/edge) voice assistant systems β define the edge/cloud allocation model, routing criteria, and the port/adapter boundaries that allow functions to move between cloud and vehicle without redesign.Design and oversee end-to-end prototype pipelines (ASR β LLM β TTS), including integration into early vehicle platforms (EΒ³, SDV) and test operation in lab and driving environments.Drive evaluation and benchmarking of base components: on-device ASR for edge hardware, modern TTS models (latency, robustness, audio quality, energy efficiency), and candidate LLMs for dialogue-based assistant functions via Azure AI Foundry.Define RAG architectures for improved knowledge coverage and robustness, including embedding strategies and small-scale vector stores for specific automotive use cases.Architect the embedded/edge AI optimisation workstream: quantisation, pruning and distillation for ARM/DSP/NPU targets, with explicit attention to memory footprint, energy consumption, wake-word efficiency and thermal behaviour.Shape the research agenda for future key technologies β continual learning and on-device model adaptation (anti-drift), multimodal interaction (speech + image + vehicle sensor data), emotion and sentiment recognition, privacy-compliant on-device personalisation and long-term memory, multi-speaker handling, speaker identification and anti-spoofing.Embed safety, security and privacy mechanisms into concepts from the start (differential privacy, secure enclaves, automotive security standards) and ensure interoperability with vehicle architectures (SOA, microservices, zonal architecture).Apply and enforce hexagonal architecture as the structural standard, so prototypes are transferable into series development rather than thrown away.Produce the technical concepts, architecture decision records and evaluation reports that hand research results over to the series development work package β and defend them in review with the client's architects.Lead the offshore engineering team as Location Lead: technical direction for AI/ML and AI Ops engineers, code and design review standards, work breakdown and estimation across a high-throughput ticket flow (~30 tickets/sprint, sizes S/M/L), onboarding and skills growth.Act as the offshore technical counterpart to the onshore solution architect and client engineering teams; run technical alignment across time zones and represent the location in architecture boards and sprint ceremonies.
Skills
Must have
Edge / on-device AI: model optimisation (quantisation, pruning, distillation) and deployment to constrained hardware β ARM, DSP, NPU, or mobile/embedded equivalents.Expert Python.Hands-on with LLM frameworks in production: LangChain, LangGraph or equivalent.3+ years as architect or technical lead, with end-to-end ownership of a system's design (within 8+ years total engineering experience).Track record of taking AI/ML or LLM systems from prototype into production β not research-only.Practical experience building RAG systems: embedding models, vector stores, retrieval evaluation.Azure OpenAI Service or OpenAI API in practice, plus Azure cloud (AKS or Container Apps).Integration experience with at least one ASR or TTS engine (Whisper, Azure Speech, Cerence or equivalent).Experience leading a distributed engineering team (5β10 engineers) in Scrum/Kanban.
Nice to have Skills
Hexagonal / ports-and-adapters architecture and API design (REST, gRPC, protobuf, OAuth2) β the client mandates this architecture style, so the candidate must be willing to adopt it quicklyAutomotive / in-vehicle software context: infotainment platforms (MIB3, EΒ³, SDV), Android Automotive / AAOS, AIDL, Viwi, vehicle telemetry services, SOA or zonal E/E architectures.In-vehicle speech processing specifics: noise cancellation, far-field microphones, wake-word engines, barge-in, multi-speaker cabin scenarios.Edge inference toolchains: ONNX Runtime, TensorRT, TFLite, OpenVINO; audio signal processing with pyAudio / librosa / pydub.Observability and evaluation for LLM systems: LangFuse, OpenTelemetry, and evaluation frameworks (DeepEval, Ragas, PromptFlow); quality metrics such as BLEU, WER, CER, faithfulness and hallucination rate.Commercial automotive speech/emotion stacks (Cerence, HumeAI) and Google TTS.Privacy-preserving ML: differential privacy, federated/aggregate learning signals, secure enclaves, on-device personalisation.Speaker identification and anti-spoofing.Data/persistence breadth: PostgreSQL + pgvector, MongoDB, CosmosDB; IaC with Terraform/Terragrunt.Awareness of automotive quality and security frameworks (A-SPICE, ISO/SAE 21434, TISAX) and of EU AI Act implications for AI systems.German language skills.
Languages:
English - C1 Fluent