Data Engineer

Sojo Data — Nepal · Posted ~6 hours ago

Skills

Python Apache Spark Batch data pipelines Streaming data pipelines Data processing APIs Amazon S3 CI/CD Data validation Data testing Data monitoring Data quality

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Develop and maintain scalable batch and streaming data pipelines using Python and Apache Spark. You will integrate data with external systems and cloud services, implement CI/CD and data-quality controls, and collaborate across teams to deliver reliable, production-ready data solutions.

Highlights

Build scalable batch and streaming data pipelines while focusing on performance, reliability, cost efficiency, testing, data quality, and modern AI-assisted engineering practices.

Description

Data EngineerResponsibilitiesDesign, develop, and maintain scalable batch and streaming data pipelines using Python and Apache Spark.Develop and optimize data processing solutions with a focus on performance, scalability, reliability, and cost efficiency.Write clean, modular, reusable, well-tested, and production-ready code.Integrate data platforms with external systems through APIs, cloud-based services, and data sources such as Amazon S3.Implement and maintain CI/CD workflows for data engineering projects.Apply appropriate data validation, testing, monitoring, and data quality practices.Collaborate with engineers, stakeholders, and cross-functional teams to understand business and technical requirements and deliver effective data solutions.Use modern AI-assisted engineering and development tools to improve productivity, code quality, and problem-solving.Understand business requirements and consider the broader impact of data engineering and architectural decisions.Contribute to continuous improvement of data platforms, pipelines, development standards, and engineering practices.RequirementsMust-HaveStrong hands-on experience with Apache Spark / Spark SQL.Strong Python programming skills with experience developing clean, modular, maintainable, and production-ready code.Experience designing and implementing batch and streaming data pipelines.Experience with CI/CD workflows for data engineering projects.Knowledge of data pipeline optimization, performance tuning, scalability, and cost management.Experience with data QA, validation, testing, monitoring, and data quality practices.Experience working with APIs and cloud-based data sources, such as Amazon S3.Strong problem-solving, communication, collaboration, and ownership skills.Ability to understand business requirements and translate them into reliable and scalable data solutions.Preferred / Nice to HaveHands-on experience with the Databricks Lakehouse Platform is preferred but not mandatory.Knowledge of Delta Lake, Delta Tables, Unity Catalog, DLT/Lakeflow, Streaming Tables, Materialized Views, and Checkpoints is a plus.Experience with Databricks Asset Bundles (DAB) and YAML-based deployment configurations is a plus.Experience with other cloud data platforms, data lakes, lakehouses, or distributed data processing technologies is also valuable.Familiarity with AI-assisted development tools, agentic coding, AI plugins, and AI-powered engineering workflows.Understanding of frontend and backend systems and how application events generate, process, and consume data.Candidates with strong Python + Apache Spark + data engineering experience are encouraged to apply even if they do not have prior Databricks experience.Technology StackCore: Python, Apache Spark, Spark SQL, SQL Cloud/Data: Amazon S3, APIs, Cloud Data Platforms Engineering: Git, CI/CD, Testing, Data Quality Preferred: Databricks, Delta Lake, Unity Catalog, DLT/Lakeflow, DAB