Senior Multimodal Data & AI Infrastructure Expert

Sparagus — Netherlands · Posted ~7 hours ago

Senior Full-time Onsite

Skills

data infrastructure AI infrastructure multimodal computing distributed systems resource scheduling vector storage and retrieval distributed caching unstructured data processing large language models AI agents system architecture vector storage AI4DB Data Agents LLMs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead the architecture and development of next-generation multimodal data infrastructure for large-scale AI workloads. You will work across heterogeneous computing, distributed storage and caching, vector retrieval, scheduling, and intelligent autonomous management to enable advanced language-model and AI-agent ecosystems.

Highlights

Shape next-generation AI data infrastructure at massive scale, working across multimodal computing, distributed systems, storage, retrieval, and autonomous data management.

Description

Location: Amsterdam, the Netherlands Employment Type: Full-time, Permanent About the Role We are looking for a senior Multimodal Data & AI Infrastructure Expert to help shape the next generation of data infrastructure for the AI Agent era. You will drive the architecture and core technology development of a next-generation multimodal intelligent data platform, working across heterogeneous computing, multimodal computing engines, vector storage and retrieval, distributed caching, AI4DB, and Data Agent technologies. The role focuses on building end-to-end capabilities for massive-scale heterogeneous and unstructured data — spanning resource scheduling, computation, retrieval, storage, and intelligent autonomous management. You will work at the intersection of Data Infrastructure and AI, developing core technologies optimized for large language models, multimodal workloads, and AI Agent ecosystems. What You Will Do Multimodal Data Infrastructure Design and develop a unified scheduling and management platform for heterogeneous computing resources, including CPUs, GPUs, and NPUsDevelop serverless resource pooling and management capabilities across multiple computing enginesOptimize heterogeneous resource scheduling, system performance, and overall resource utilizationMultimodal Computing Engines Develop and optimize computing engines for large-scale unstructured and multimodal dataEnable efficient processing and analysis of text, images, video, and other data typesDesign next-generation query optimizers and hybrid execution enginesSupport high-performance retrieval and processing of vector, textual, geospatial, and other dataVector Storage & Retrieval Design and develop systems for multimodal vector retrieval, storage, metadata management, and access controlDevelop intelligent storage optimization and index lifecycle management capabilitiesIntegrate data infrastructure with relevant open-source ecosystemsDistributed Caching Build high-performance distributed caching services for multimodal data platformsDevelop high-speed near-compute caching capabilities for computing enginesOptimize efficient east-west data transfer across distributed computing enginesAI4DB & Data Agents Explore and apply AI4DB (AI for Databases) and LLM Agent technologiesDevelop intelligent Data Agents and Skills for next-generation data platformsEnable greater automation and intelligence across workload development, data storage, data analysis, operations, and system maintenance Required Qualifications Strong programming skills in languages such as C, C++, Python, or JavaStrong R&D background and technical expertise in one or more of the following areas:Database SystemsBig Data SystemsDistributed SystemsHigh-Performance ComputingStrong understanding of heterogeneous computing architectures and performance considerations involving CPUs, GPUs, and NPUsHands-on experience with heterogeneous resource management and schedulingExperience in cloud computing platform development and maintenanceExperience with DevOps-related engineering practices Preferred Qualifications Experience in one or more of the following would be highly relevant: Query optimizersExecution enginesStorage enginesDistributed storage systemsVector databases, vector indexing, or high-performance retrieval systemsLarge-scale multimodal or unstructured data processingLLM fine-tuning and reinforcement learningNLP, computer vision, or time-series data processingIntegration of AI technologies with database or data systemsAI4DBLLM Agents / Data AgentsHigh-performance distributed cachingServerless and heterogeneous computing infrastructure What We Offer A highly competitive compensation packageRelocation allowance for candidates relocating for the positionAnnual performance bonusAttractive short-term and long-term incentive programsThe opportunity to work on next-generation data infrastructure at the intersection of AI, database systems, distributed computing, and multimodal technologies Who Should Apply? We are particularly interested in experienced database, distributed systems, data infrastructure, and high-performance computing experts who have worked on core system technologies rather than only application-level data engineering. If you have built database engines, distributed data systems, multimodal computing infrastructure, vector retrieval systems, or AI-native data platforms and are interested in building the next generation of infrastructure for LLM and AI Agent ecosystems, we would be very happy to connect.