Senior Big Data Developer

Themesoft Inc. — Canada · Posted ~22 hours ago

Senior Full-time Hybrid

Skills

Big Data Spark Spark SQL Scala Unix/Shell scripting Hadoop ETL/ELT Data pipelines HiveQL SQL HDFS API integration Cloudera CDP Unix/Shell Parquet ORC Avro

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a senior data engineering team building large-scale data pipelines and backend systems in an enterprise environment. You will develop Spark and Scala applications, optimize ETL/ELT workloads, write advanced SQL queries, work with distributed storage formats, tune performance, and modernize legacy data-processing applications.

Highlights

Hybrid senior engineering role focused on large-scale data systems, with opportunities to design and optimize sophisticated pipelines, work with modern distributed data platforms, and solve challenging performance and migration problems.

Description

Position: Big Data Developer Location: Toronto – Hybrid Skills: Big Data & SparkSQLUnix/Shell Scripting Role Overview We are looking for a Senior Backend Developer with 5+ years of experience in big data engineering, API integration, and AI-assisted development. The ideal candidate will design, build, and maintain scalable data pipelines and backend systems in a enterprise environment. Key Responsibilities Big Data & Spark Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clustersBuild and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQLTune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)Migrate Spark 2 applications to Spark 3 on Cloudera CDP platformsWork with Parquet, ORC, Avro file formats on HDFS SQL & Data Engineering Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregationsDesign and maintain Hive external/managed tables and partitioned datasetsOptimize slow-running queries and resolve correlated subquery issuesWork with HDFS encryption zones and data governance requirements Unix / Shell Scripting Develop and maintain bash shell scripts for job orchestration and automationHandle error management, return codes, logging and alerting in shell scriptsManage HDFS operations (hdfs dfs commands), file transfers, and data validationManage Kerberos authentication (kinit, keytab handling) API Extraction & Integration Build scripts and pipelines to extract data from REST APIs using curl and PythonHandle OAuth2 token generation, bearer token refresh and API health checksParse and process JSON API responses and load into HDFS/HiveManage pagination, error handling and retry logic for API callsWork with enterprise API gateways and URL parameter construction AI & Copilot Capabilities Leverage GitHub Copilot / AI coding assistants to accelerate developmentUse AI tools for code review, SQL generation, script debugging and documentationContribute to AI-assisted data quality and anomaly detection pipelinesExplore and implement LLM-based automation for repetitive data engineering tasks Scheduling & Orchestration Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cronBuild and maintain Ansible playbooks for automated deploymentsManage deployment pipelines including artifact versioning, Vault secret injection and environment-specific configurationMonitor job health, handle failures and implement alerting Nice to Have Experience with Cloudera CDP (7.x) and migration from HDPKnowledge of Kerberos, Vault, HDFS encryption zonesFamiliarity with CI/CD pipelines (Helios, GitHub Actions)Experience with MSSQL / JDBC connectivity from SparkUnderstanding of AML / Financial regulatory data domains Regards Patrick Fernandez Talent Acquisition Group - Strategic Recruitment Manager