Summary
✨ AI‑Generated
Join a senior data engineering team building large-scale data pipelines and backend systems in an enterprise environment. You will develop Spark and Scala applications, optimize ETL/ELT workloads, write advanced SQL queries, work with distributed storage formats, tune performance, and modernize legacy data-processing applications.
Highlights
Hybrid senior engineering role focused on large-scale data systems, with opportunities to design and optimize sophisticated pipelines, work with modern distributed data platforms, and solve challenging performance and migration problems.
Description
Position: Big Data Developer
Location: Toronto – Hybrid
Skills:
Big Data & SparkSQLUnix/Shell Scripting
Role Overview
We are looking for a Senior Backend Developer with 5+ years of experience in big data engineering, API integration, and AI-assisted development.
The ideal candidate will design, build, and maintain scalable data pipelines and backend systems in a enterprise environment.
Key Responsibilities
Big Data & Spark
Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clustersBuild and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQLTune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)Migrate Spark 2 applications to Spark 3 on Cloudera CDP platformsWork with Parquet, ORC, Avro file formats on HDFS
SQL & Data Engineering
Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregationsDesign and maintain Hive external/managed tables and partitioned datasetsOptimize slow-running queries and resolve correlated subquery issuesWork with HDFS encryption zones and data governance requirements
Unix / Shell Scripting
Develop and maintain bash shell scripts for job orchestration and automationHandle error management, return codes, logging and alerting in shell scriptsManage HDFS operations (hdfs dfs commands), file transfers, and data validationManage Kerberos authentication (kinit, keytab handling)
API Extraction & Integration
Build scripts and pipelines to extract data from REST APIs using curl and PythonHandle OAuth2 token generation, bearer token refresh and API health checksParse and process JSON API responses and load into HDFS/HiveManage pagination, error handling and retry logic for API callsWork with enterprise API gateways and URL parameter construction
AI & Copilot Capabilities
Leverage GitHub Copilot / AI coding assistants to accelerate developmentUse AI tools for code review, SQL generation, script debugging and documentationContribute to AI-assisted data quality and anomaly detection pipelinesExplore and implement LLM-based automation for repetitive data engineering tasks
Scheduling & Orchestration
Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cronBuild and maintain Ansible playbooks for automated deploymentsManage deployment pipelines including artifact versioning, Vault secret injection and environment-specific configurationMonitor job health, handle failures and implement alerting
Nice to Have
Experience with Cloudera CDP (7.x) and migration from HDPKnowledge of Kerberos, Vault, HDFS encryption zonesFamiliarity with CI/CD pipelines (Helios, GitHub Actions)Experience with MSSQL / JDBC connectivity from SparkUnderstanding of AML / Financial regulatory data domains
Regards
Patrick Fernandez
Talent Acquisition Group - Strategic Recruitment Manager