Summary
✨ AI‑Generated
A senior backend and big data engineering role focused on designing scalable data pipelines and high-volume processing systems. You will work with Spark, Scala, SQL, distributed storage, enterprise data platforms, and performance optimization in a large-scale environment.
Highlights
Senior big data engineering role with hybrid work, focused on large-scale Spark processing, scalable data pipelines, performance optimization, and enterprise data platforms.
Description
Role : Big Data Developer
Location : Toronto, ON (Hybrid- 3/4 days onsite)
Top 3 skills required for this role:
Big Data & SparkSQLUnix/Shell Scripting
Role Overview
We are looking for a Senior Backend Developer with 5+ years of experience in big data engineering, API integration, and AI-assisted development.
The ideal candidate will design, build, and maintain scalable data pipelines and backend systems in a enterprise environment.
Key Responsibilities
Big Data & Spark
Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clustersBuild and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQLTune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)Migrate Spark 2 applications to Spark 3 on Cloudera CDP platformsWork with Parquet, ORC, Avro file formats on HDFSSQL & Data Engineering
Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregationsDesign and maintain Hive external/managed tables and partitioned datasetsOptimize slow-running queries and resolve correlated subquery issuesWork with HDFS encryption zones and data governance requirementsUnix / Shell Scripting
Develop and maintain bash shell scripts for job orchestration and automationHandle error management, return codes, logging and alerting in shell scriptsManage HDFS operations (hdfs dfs commands), file transfers, and data validationManage Kerberos authentication (kinit, keytab handling)API Extraction & Integration
Build scripts and pipelines to extract data from REST APIs using curl and PythonHandle OAuth2 token generation, bearer token refresh and API health checksParse and process JSON API responses and load into HDFS/HiveManage pagination, error handling and retry logic for API callsWork with enterprise API gateways and URL parameter constructionAI & Copilot Capabilities
Leverage GitHub Copilot / AI coding assistants to accelerate developmentUse AI tools for code review, SQL generation, script debugging and documentationContribute to AI-assisted data quality and anomaly detection pipelinesExplore and implement LLM-based automation for repetitive data engineering tasksScheduling & Orchestration
Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cronBuild and maintain Ansible playbooks for automated deploymentsManage deployment pipelines including artifact versioning, Vault secret injection and environment-specific configurationMonitor job health, handle failures and implement alerting
Nice to Have
Experience with Cloudera CDP (7.x) and migration from HDPKnowledge of Kerberos, Vault, HDFS encryption zonesFamiliarity with CI/CD pipelines (Helios, GitHub Actions)Experience with MSSQL / JDBC connectivity from SparkUnderstanding of AML / Financial regulatory data domains.
We are an Equal Opportunity Employer.
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, citizenship status, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law.
https://www.e-verify.gov/sites/default/files/everify/posters/IER_RighttoWorkPoster.pdf