Description
Databricks Developer
1.
PURPOSE
The Cloud Data Engineer (Azure Databricks) in the Healthcare domain is responsible for building, testing, and maintaining data pipelines on Azure Databricks to ensure reliable, accessible, high-quality data across the organization.
This role supports clinical, operational, and analytical teams by transforming raw data into actionable insights, helping drive informed decision-making and improved patient outcomes.
Working under the guidance of senior data engineering and architecture leadership, this role contributes to the design and ongoing development of the Enterprise Data Warehouse (EDW), with a primary focus on Databricks-based ingestion, transformation, and data Lakehouse development.
2.
TECHNICAL SKILLS
Hands-on experience building and maintaining ETL/ELT pipelines on Azure Databricks using PySpark, Spark SQL, and Databricks notebooks for batch and streaming data processing.Working knowledge of Delta Lake concepts, including Delta tables, schema evolution, time travel, and optimizing table performance (OPTIMIZE, Z-ORDER, VACUUM).Experience developing and orchestrating data workflows using Databricks Workflows/Jobs and integrating them with Azure Data Factory (ADF) for end-to-end pipeline orchestration.Familiarity with the Medallion architecture (bronze/silver/gold layers) and building curated datasets for downstream analytics and reporting.Exposure to Unity Catalog or similar governance tooling for managing data access, lineage, and cataloging within a Databricks/Azure environment.Solid working knowledge of SQL for querying, transforming, and validating data within Azure SQL Database, Azure Synapse Analytics, and Databricks SQL.Experience working with Azure Data Lake Storage (ADLS Gen2) and Azure Blob Storage to manage structured and unstructured data assets.Ability to build pipelines that ingest data from databases, SFTP, file-based systems, and REST APIs into the Databricks/Azure ecosystem.Working knowledge of data modeling fundamentals, including conceptual, logical, and physical model concepts for transactional and analytical systems.Basic understanding of the broader Azure data and AI ecosystem (Azure Synapse Analytics, Azure ML, Azure SQL Database) and how Databricks pipelines integrate with AI/ML workflows.Exposure to operationalizing AI/ML pipelines, such as preparing training datasets and integrating model outputs into data workflows using Databricks or Azure ML.Familiarity with data normalization/denormalization principles and relational database design; exposure to data modeling tools (ER/Studio, Erwin, or similar) is a plus.Ability to follow and help uphold coding standards, participate in code reviews, and support data validation and testing (unit, integration, and regression).Working understanding of data standards, naming conventions, and governance practices for enterprise and clinical datasets.Exposure to healthcare data vocabularies such as SNOMED CT, LOINC, ICD-10, and RxNorm is a plus, but not required.Working knowledge of CI/CD practices, including version control (e.g., Git), branching strategies, and automated deployment basics.Exposure to data quality practices, including data profiling, validation, and basic cleansing strategies.Basic understanding of metadata management and data lineage concepts that support auditability, compliance, and governance.Ability to collaborate effectively with business, clinical, and technical stakeholders to support data initiatives aligned with organizational goals and regulatory requirements (e.g., HIPAA, GDPR).Ability to manage individual task timelines, prioritize work, and adapt to changing requirements to deliver quality data solutions on time.3.
REQUIREMENTS
2β4 years of experience designing and implementing conceptual, logical, and physical data models in transactional and/or analytical environments; healthcare domain experience is a plus.2β4 years of hands-on experience building cloud data engineering solutions; prior exposure to the healthcare domain is a plus but not required.3β5 years of experience with data normalization and denormalization principles, relational database design, and basic performance optimization.3β5 years of experience developing cloud data pipelines (data warehousing, data lake, or data mart work) using ETL/ELT tools and database management systems.2β3 years of experience translating business and/or clinical requirements into data structures that support reporting and analytics.Exposure to healthcare industry-standard data models (e.g., CDISC SDTM/ADaM, OMOP, FHIR, HL7) is a plus but not required.2β4 years of experience developing and optimizing data pipelines using Azure Databricks (PySpark/Spark SQL), with working exposure to Azure Data Factory and/or Azure Synapse Analytics.2β4 years of experience implementing data models using Azure Data Lake Storage Gen2 and Delta Lake to support analytics, BI, and reporting use cases.2β3 years of experience developing and maintaining ADF pipelines or Databricks Workflows, supporting reliable data ingestion into Azure Data Lake, Azure SQL, or Synapse Analytics.1β2 years of exposure to data quality, metadata management, or data lineage practices, business glossaries, or documentation supporting model traceability.2β3 years of experience building Databricks or ADF pipelines to orchestrate data ingestion, transformation, and loading (ETL/ELT) across sources, including SQL Server, REST APIs, and Blob Storage.Exposure to AI/ML pipeline support (e.g., preparing training datasets) is a plus, but not required for this level.4.
INTERPERSONAL SKILLS
Good communication (written and verbal), listening, and interpersonal skills; able to build rapport with team members and stakeholders.Demonstrated ability to collaborate effectively within a team and across departments in a heterogeneous environment.Willingness to build working relationships with internal customers, take ownership of assigned issues, and escalate appropriately.Ability to communicate technical work clearly to both technical and non-technical stakeholders.Ability to participate in requirements-gathering discussions and translate defined requirements into technical implementation tasks.Openness to receiving mentorship and coaching, with growing ability to support and guide junior team members over time.Benefits and Schedules
Salary: 1,800,000 NPR β 2,400,000 NPR (Annual)
Working Hours: 10:00 PM-6:00 AM (Onsite)
Location: Kathmandu
Contact: contactus@stataAI.com, #: 9848584325