Summary
✨ AI‑Generated
Join an enterprise data engineering team responsible for designing, optimizing, and maintaining scalable data integration and processing solutions. You will build pipelines with Informatica and Spark, develop robust ETL/ELT workflows, process large structured and semi-structured datasets, and work with distributed data technologies while addressing performance and data-quality challenges.
Highlights
Data engineering role with substantial hands-on responsibility for enterprise-scale data integration and processing. The position offers work with large datasets, modern distributed data technologies, reusable data components, and both batch and near-real-time processing workflows.
Description
We are looking for a Data Engineer with 4–7 years of experience, having strong hands-on expertise in Informatica Big Data Management (BDM) and Apache Spark.
The candidate will be responsible for developing, optimizing, and maintaining scalable data integration and processing solutions across enterprise data platforms.
Responsibilities:
Design, develop, and maintain data pipelines using Informatica BDM.Develop and optimize large-scale data processing workflows using Apache Spark/PySpark.Build robust ETL/ELT processes for batch and near-real-time data processing.Perform data extraction, transformation, cleansing, and loading from multiple source systems.Develop Informatica mappings, workflows, transformations, and reusable components.Work with Hadoop ecosystem technologies such as HDFS, Hive, and related components.Write and optimize Spark jobs for processing large volumes of structured and semi-structured data.Troubleshoot performance, data quality, and production issues across Informatica and Spark pipelines.Optimize ETL workflows, Spark jobs, SQL queries, and data processing performance.Collaborate with Data Architects, Data Analysts, and other engineering teams to implement data solutions.Implement data validation, reconciliation, error handling, and monitoring mechanisms.Participate in deployment, production support, and release activities.Follow data engineering best practices for scalability, reliability, security, and maintainability.Prepare technical documentation for data pipelines, mappings, workflows, and operational processes.
Requirements:
4–7 years of hands-on experience in Data Engineering, ETL, or Big Data technologies.Strong hands-on experience with Informatica Big Data Management (BDM).Strong experience with Apache Spark; PySpark experience is highly preferred.Good understanding of Informatica mappings, workflows, transformations, and data integration processes.Strong SQL skills with experience in query development and performance optimization.Experience working with Hadoop ecosystem technologies such as HDFS and Hive.Experience handling large-scale structured and unstructured datasets.Strong understanding of ETL/ELT concepts and data warehousing principles.Experience with batch data processing and distributed data processing.Good knowledge of Python/Scala for data processing and automation.Experience with Linux/Unix environments and shell scripting.Strong troubleshooting and performance-tuning skills.Understanding of data quality, data validation, and error-handling practices.Experience with version control tools such as Git.Good communication, analytical, and problem-solving skills.Ability to work independently as well as collaboratively in a cross-functional team.