Python Data Engineer

Apptoza Inc — Canada · Posted ~2 hours ago

Contract Hybrid

Skills

Python Pandas Polars Docker Kubernetes Data pipelines ClickHouse NATS Pytest Distributed computing Data processing

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A data engineering role based in a hybrid work environment, focused on designing and maintaining high-performance Python data pipelines. The position involves large datasets, containerized applications, distributed processing, analytical databases, event-driven architectures, and comprehensive automated testing.

Highlights

Build high-performance data pipelines, work with large-scale datasets and distributed systems, and develop scalable containerized solutions in a data engineering environment.

Description

Python Developer Toronto, ON - Hybrid (4 Days WFO) 6-12 months Role Descriptions: Job Description We are seeking a skilled Python Developer to join our data engineering team. You will design| develop| and maintain high-performance data processing pipelines using modern Python frameworks and tools. In this role| youll work with large-scale datasets| containerized systems| and distributed computing platforms to deliver robust data solutions. Key Responsibilities Develop and optimize data manipulation workflows using pandas and polars to handle large datasets efficiently. Design and implement containerized applications using Docker and Kubernetes to ensure scalable| reliable deployments. Build and maintain data pipelines integrating with Click House columnar databases for analytical workloads. Develop event-driven architectures using NATS messaging systems for asynchronous data processing. Write comprehensive unit tests using pytest to ensure code quality and reliability. Implement distributed computing solutions with Dask for processing data beyond single-machine memory constraints. Manage version control using Git and collaborate on code repositories following best practices. Required Skills and Experience Python & Data Processing: Advanced proficiency in pandas and polars for data manipulation| transformation| and analysis. Experience optimizing code performance for large datasets. Containerization & Orchestration: Hands-on experience with Docker for building container images and composing multi-container applications. Knowledge of Kubernetes for container orchestration and deployment management. Data Infrastructure: Working knowledge of ClickHouse or similar columnar databases for OLAP workloads and analytical queries. Messaging & Streaming: Familiarity with NATS.io for building message-driven systems and asynchronous workflows. Testing & Quality Assurance: Proficiency with pytest for writing unit tests| integration tests| and maintaining code coverage standards. Distributed Computing: Experience with Dask for parallel processing and handling out-of-core computations. Version Control: Strong command of Git workflows| branching strategies| and collaborative development practices. Preferred Qualifications Experience with additional Python libraries for data science and machine learning. Familiarity with CI/CD pipelines and DevOps practices. Background in financial services or capital markets data systems.