Lead Data Engineer

Megthink Solutions India Pvt Limited — United States · Posted ~2 hours ago

Lead Full-time Hybrid No Visa

Skills

Data platform architecture Lakehouse architecture Azure Data Lake Storage Gen2 Delta Lake Metadata-driven ingestion frameworks Batch and streaming data ingestion Change Data Capture (CDC) PySpark Scala Spark Structured Streaming Azure Event Hubs Apache Kafka Databricks Azure Data Factory Azure DevOps CI/CD Terraform Databricks Asset Bundles Azure Monitor Log Analytics Code reviews Production incident debugging Spark cluster tuning Data schema evolution Data retention and partitioning AI-assisted data mapping Software testing Python Azure ADLS Gen2 Apache Spark AI-assisted data engineering

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An established technology organization is seeking a hands-on Lead Data Engineer to architect and operate enterprise-scale data platforms. You will design lakehouse layers, build metadata-driven ingestion frameworks, optimize distributed processing, establish reliable deployment pipelines, and introduce AI-assisted data mapping. The role suits an experienced technical leader who combines architectural judgment with strong implementation, testing, and production troubleshooting skills. Applicants must be eligible to work as US citizens, and the position requires an onsite or hybrid working arrangement.

Highlights

Lead a hands-on data engineering function with ownership of modern lakehouse architecture, scalable ingestion frameworks, and production-grade data pipelines. Work with advanced cloud and distributed processing technologies, improve engineering standards, and introduce AI-assisted automation in a collaborative technical environment.

Description

Role: Lead Data Engineer (Hands-On) Location: Cary, NC (On-site / Hybrid) Experience: 12-18 years Employment: Full-Time Eligibility: US Citizens only What the Role Owns:- Data platform architecture and engineering: lakehouse architecture (Bronze / Silver / Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, retention).Metadata-driven, parameterized ingestion frameworks for batch files, database extracts, CDC feeds and streaming (Azure Event Hubs / Kafka, Spark Structured Streaming).Canonical PySpark and Scala Spark jobs, coding and testing standards, PR reviews, production incident debugging, Spark cluster tuning and cost guardrails.CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset Bundles and Terraform; observability with Azure Monitor and Log Analytics.AI-augmented ingestion and canonical mapping: auto-generated bridgedocuments, DML, canonical table definitions; AI-assisted source-to-canonical mapping with human review gate.AI-driven data quality, anomaly detection (data drift, schema drift, volume shifts, reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data.Semantic layer and knowledge graph, plus a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security.Governance and leadership: Unity Catalog (lineage, access control, PII standards), Architecture Review Boards and AI governance forums, mentoring engineers, documentation. Must have Skills & Experience:- Expert-level Python, Scala and PySpark: production-ready, modular, well-tested solutions; Spark workload troubleshooting; optimizing large-scale batch and streaming pipelines using Delta Lake.Strong SQL and data modelling (dimensional and normalised), schema design, data contracts.Databricks expertise: Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.Azure data stack: ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks), Azure Event Hubs.3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, prompt engineering.Evaluation discipline: golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback.Hands-on with LangChain, LlamaIndex or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).Metadata-driven frameworks: schema inference, data profiling, lineage, catalogs.12-18 years of total experience in data engineering / data platform delivery.Proven enterprise-scale delivery of a medallion / lakehouse architecture.Azure security and governance: Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, PII handling.CI/CD and IaC: Azure DevOps, Terraform, Databricks Asset Bundles, automated testing of data pipelines.Clear technical writing and ability to present and defend designs to engineers and non-technical stakeholders. Strongly Preferred:- Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modelling over a lakehouse).Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale.ML-based anomaly detection on time-series or transactional financial data.Financial services or insurance domain (finance close, GL, subledger, reconciliation, actuarial data).LLMOps / MLOps: model and prompt versioning, cost governance, observability.Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.dbt, Great Expectations or similar; Workday, Prism or Accounting Center exposure.