Data Engineer

Rheinmetall — Germany · Posted ~4 hours ago

Visa History ✓

Skills

Data lakehouse architecture Data platform architecture Data storage strategy Batch data pipelines Streaming data pipelines Data ingestion Data transformation Semantic data modeling Ontology design Metadata management Partitioning Indexing Compression Archiving Machine learning data foundations AI systems Kubernetes Data lakehouse Data platforms Batch processing Stream processing Semantic data models Ontologies Machine learning AI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A data engineering role focused on designing and operating a scalable lakehouse and data platform for very large heterogeneous datasets. You will build robust batch and streaming pipelines, develop storage and access strategies for structured and unstructured data, establish semantic models and ontologies, and create a resilient foundation for machine learning, AI applications, and autonomous systems.

Highlights

Opportunity to design and implement a highly scalable data platform handling very large and diverse datasets, with a strong focus on performance, resilience, lifecycle management, and AI-ready data foundations.

Description

What We Are Looking For Design, further development, and practical implementation of a scalable data lakehouse and data platform architecture for very large, heterogeneous data setsDeveloping a long-term data and storage strategy that takes into account performance, scalability, data lifecycle, costs, availability, and operational requirementsDesigning suitable storage and access concepts for large volumes of image, video, sensor, geospatial, and other unstructured and semi-structured data, including metadata, partitioning, indexing, compression, and archivingBuilding and operating robust batch and streaming pipelines for the ingestion, processing, transformation, and delivery of data from heterogeneous sourcesCo-design and implementation of a semantic data model or ontology that enables data objects, entities, relationships, and context to be described and linked across systemsCreating a resilient data foundation for machine learning, AI-based applications, and autonomous systems, including traceable data provenance and reproducible training and validation datasetsEnsuring data lineage, data versioning, data quality, schema management, and appropriate governance and access concepts throughout the entire data lifecycleDevelopment and operation of the data platform in large, distributed Kubernetes-based infrastructures, as well as integration into Infrastructure-as-Code and CI/CD/DevSecOps processesAnalyzing and optimizing data layouts, data locality, compute proximity, and data movements for data-intensive applicationsEvaluation of suitable open-source and commercial technologies, as well as implementation ranging from proofs of concept to production-ready, scalable solutionsClose collaboration with AI/ML, software, platform, and cybersecurity teams on architectural decisions, interfaces, data products, and operational implementation What Qualifications You Should Have A degree in computer science, data engineering, software engineering, business informatics, or a comparable practical qualificationSeveral years of hands-on experience in data engineering and in building or further developing distributed data platformsA very good understanding of modern data lake/lakehouse architectures, object storage, and the fundamental trade-offs between storage, compute, data movement, performance, and costsDemonstrable hands-on experience with data pipelines, data modeling, SQL, and at least one relevant programming language, preferably PythonSolid knowledge of metadata, data quality, schema management, data lineage, and reproducible processingPractical understanding of Linux, containers, and Kubernetes, as well as the ability to independently implement, operate, and analyze data components in a distributed platformAbility to link architectural decisions to practical implementation and to systematically investigate problems down to their technical root causesA structured, team-oriented approach to work, as well as strong written and spoken English skillsExperience with very large volumes of image, video, sensor, or geodataKnowledge of ontologies, knowledge graphs, semantic data models, entity resolution, or comparable approachesExperience with technologies for distributed data processing and streaming, e.g., Spark, Flink, Trino, Kafka, or comparable solutionsExperience with open Lakehouse table formats or comparable technologies, such as Apache Iceberg, Delta Lake, or Apache HudiExperience in provisioning and versioning data for machine learning, MLOps, AI, or autonomy workloadsKnowledge of Infrastructure as Code, GitOps, or automated platform and deployment processesExperience with highly available, on-premises, edge, or intermittently disconnected data platforms A solid foundation and practical implementation experience are crucial for this role. Additional specialized knowledge is expressly not a requirement. Stärken und Erfahrungen zählen bei Rheinmetall, auch wenn vielleicht nicht alle aufgeführten Anforderungen vollständig erfüllt sind. Wir freuen uns auf Bewerber (m/w/d), die Lust haben, etwas zu bewegen und Verantwortung zu übernehmen. Wir legen Wert auf Individualität und Chancengleichheit. Schwerbehinderte Bewerber (m/w/d) werden bei gleicher Eignung besonders berücksichtigt. What We Offer You At our location in Bremen, we offer you: Company pension schemeShare purchase programme30 days of holidayAccess to corporate benefitsDeutschlandticketRelocation supportMobile workingVIVA family serviceIndividual and diverse internal and external development opportunities, including at the Rheinmetall AcademyProfessional induction process supported by digital onboarding CONTACT INFORMATION Contact Person: Ms Özge Demirkaya For questions regarding your application, please use the contact form.