Senior Software Engineer - Data and ML Platform

Jobgether — Canada · Posted ~3 hours ago

Senior Full-time

Skills

backend engineering cloud infrastructure data engineering ML operations production systems Cloud Data Engineering Machine Learning

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior engineering position focused on building and operating data and machine learning platforms. The role combines software engineering, cloud infrastructure, and production reliability improvements.

Highlights

High autonomy role with ownership of production systems, architecture decisions, and collaboration with technical teams.

Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer – Data & ML Platform based in Canada. This is a hands-on, high-ownership engineering role focused on the production systems and internal data platform supporting machine learning and operations research. You will work at the intersection of backend engineering, cloud infrastructure, data engineering, and ML operations. You’ll own and evolve live production systems while improving their reliability, scalability, observability, and maintainability. The role provides significant autonomy to shape cloud architecture, infrastructure, deployment practices, and internal developer tooling. You’ll partner closely with data scientists, operations research specialists, product teams, and engineers to bring research from experimentation into production. As the platform grows, you’ll help define the technical roadmap for data, ML infrastructure, and engineering practices. This is an opportunity to make a direct impact in a small, collaborative team where strong technical ownership and pragmatic problem-solving are highly valued. Accountabilities Own, operate, and continuously improve backend services running in Azure, including serverless applications, batch workloads, and ML inference endpoints. Manage deployments and production reliability across environments, including CI/CD, monitoring, alerting, incident response, and operational runbooks. Optimize cloud compute environments through effective management of autoscaling, containers, identity, resources, and costs. Improve the reliability, scalability, performance, and maintainability of existing production systems through incremental, pragmatic improvements. Design and build reliable ETL/ELT pipelines that transform relational and document-based data into analysis-ready datasets. Contribute to the development of a lakehouse-style analytical layer and the supporting cloud infrastructure. Build and maintain infrastructure as code using technologies such as Terraform or Bicep across multiple environments. Manage cloud services including storage, application hosting, identity and access management, Key Vault, and cost optimization. Implement data-quality controls, monitoring, logging, and freshness alerts across data workflows while ensuring appropriate security and access controls. Develop internal services, APIs, and developer tooling as platform and team requirements evolve. Build infrastructure and tooling that enable ML and Operations Research specialists to experiment, deploy, evaluate, and reproduce their work reliably. Convert research prototypes into production-ready services and workflows, including model packaging, versioning, deployment, and continuous integration. Build automated evaluation and benchmarking pipelines to monitor model performance, drift, and system reliability. Partner with Data Science and Operations Research specialists to operationalize experiments and integrate optimization and ML solutions into production products. Establish strong engineering practices around code quality, automated testing, version control, CI/CD, peer reviews, and maintainability. Collaborate with Product, Data Science, Operations Research, and Engineering teams on architecture, scalability, reliability, performance, and pragmatic technical solutions. Requirements Bachelor’s or Master’s degree in Computer Science, Software Engineering, Data Engineering, or a related field, or equivalent practical experience. Strong programming skills in Python and SQL. Strong understanding of APIs, backend service design, distributed systems, and production software engineering. Experience building and operating production data pipelines end to end, including retries, idempotency, backfills, orchestration, and data-freshness monitoring. Hands-on experience designing and operating production services in Azure or another major cloud platform, with the ability and interest to develop deep expertise in Azure. Practical understanding of cloud infrastructure, including serverless and batch compute, object storage, identity and access management, monitoring, and resource management. Experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM. Strong knowledge of Git, CI/CD, automated testing, and modern software development practices. Experience working with ML or Operations Research codebases and model artifacts, including the ability to read, run, package, and deploy models and research workflows. Proven experience owning live production systems and confidently taking over an existing codebase, understanding its architecture, and improving it over time. Ability to work autonomously and make pragmatic technical decisions in a small, evolving engineering environment. Strong communication and collaboration skills, with the ability to translate technical and business needs into practical engineering solutions. Experience with Delta Lake, Parquet, lakehouse architectures, DuckDB, Polars, or dbt is an advantage. Familiarity with Azure Machine Learning, MLflow, DVC, Durable Functions, Airflow, Dagster, or Prefect is beneficial. Experience with Kubernetes, distributed workloads, Azure data governance and security, DevOps, SRE, observability, or reliability engineering is a plus. Background in optimization, logistics, transportation, or large-scale ML systems is considered an asset. Benefits Competitive health benefits and paid time off. Remote work opportunity in Canada. Equity opportunity after the first year. Opportunities for professional growth, learning, and career advancement. High degree of autonomy and the opportunity to see your technical ideas implemented. Hands-on exposure to cloud infrastructure, data platforms, ML systems, and Operations Research. Opportunity to work closely with a tight-knit, collaborative engineering and technical team. A culture that encourages knowledge sharing, innovation, ownership, and celebrating team achievements. How Jobgether Works We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.