Data Engineer, Sr. (WFH)

Torentify Official — United States · Posted ~2 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

About the Company Transflo provides mobile, telematics, and business process automation software for the transportation and logistics industry. Its SaaS and AI solutions help freight carriers, brokers, and shippers streamline operations, reduce costs, and improve efficiency while supporting communication and collaboration across the supply chain. About the Role Transflo is seeking a Senior Data Engineer to lead the architecture and development of its enterprise data platform, from raw data ingestion to curated, analytics-ready data products. You’ll own key components of the data warehouse, pipeline infrastructure, and bronze-silver-gold medallion architecture used for internal analytics, operational reporting, and the company’s Data as a Service (DaaS) capabilities. This is a hands-on engineering role that combines platform architecture, data modeling, pipeline development, and performance optimization. You’ll work with a wide range of data sources, including APIs, relational and NoSQL databases, files, and streaming systems, turning them into reliable, governed datasets. The platform must support near real-time data delivery and high-volume workloads across transportation and logistics operations. Key Responsibilities •Architect, build, and maintain a scalable enterprise data warehouse on Amazon Redshift, applying sound data modeling, schema design, and performance optimization practices. •Develop and maintain bronze, silver, and gold medallion architecture layers for raw data ingestion, cleansing and standardization, and curated data products. •Design dimensional models, including star and snowflake schemas, fact and dimension tables, slowly changing dimensions (SCDs), and aggregate structures. •Optimize Redshift performance through distribution and sort key design, query plan analysis, workload management (WLM), cluster sizing, and related tuning strategies. •Build and maintain batch and streaming pipelines that ingest data from REST APIs, flat files, IBM DB2, MySQL, Amazon Aurora, Amazon DynamoDB, and PostgreSQL. •Implement real-time and near real-time streaming solutions using AWS services such as Kinesis Data Streams, Kinesis Firehose, Amazon MSK, and EventBridge. •Develop ETL/ELT workflows using AWS Glue, dbt, Apache Airflow, AWS Step Functions, or equivalent orchestration tools. •Build reliable pipelines with fault tolerance, idempotent processing, automated recovery, alerting, and observability. •Implement data quality processes for profiling, cleansing, deduplication, standardization, and validation across each layer of the platform. •Maintain data governance practices covering cataloging, lineage, metadata management, data classification, and access controls. •Define data contracts and service-level agreements for data freshness, completeness, and accuracy. •Work with data scientists, analytics engineers, BI developers, product managers, and API consumers to ensure data products meet quality and performance expectations. •Support the development of Transflo’s DaaS capabilities, enabling internal and external consumers to access curated data through governed APIs and data-sharing mechanisms. •Use Terraform and infrastructure-as-code practices to keep data infrastructure version-controlled, reproducible, and auditable. •Apply partitioning, concurrency scaling, caching, and workload management strategies to support high-volume analytical and operational workloads. •Implement platform security measures, including column-level security, row-level access controls, encryption, and audit logging. •Collaborate with software engineers, mobile platform teams, and DevOps to improve upstream data structure, documentation, and delivery. •Use AI-assisted development tools to accelerate pipeline development, automate data quality checks, and improve engineering efficiency. Required Qualifications •5+ years of professional data engineering experience building and operating production-grade data warehouses and pipeline infrastructure. •Expert-level Amazon Redshift experience, including cluster sizing, WLM configuration, distribution and sort key optimization, vacuuming, and query plan analysis. •Strong SQL skills covering complex analytical queries, window functions, common table expressions (CTEs), stored procedures, and performance tuning. •Hands-on experience integrating heterogeneous sources, including REST APIs, IBM DB2, MySQL, Amazon Aurora, Amazon DynamoDB, PostgreSQL, and file formats such as CSV, JSON, Parquet, and Avro. •Experience implementing enterprise-scale medallion architecture or equivalent layered data architectures. •Strong knowledge of dimensional modeling, star and snowflake schemas, slowly changing dimensions, and fact table granularity. •Experience building real-time or near real-time data pipelines using Amazon Kinesis, Apache Kafka, Amazon MSK, or equivalent technologies. •Proficiency with ETL/ELT orchestration tools such as AWS Glue, dbt, Apache Airflow, or AWS Step Functions. •Experience with data governance practices, including data catalogs, lineage tracking, metadata tagging, and access control frameworks. •Hands-on Terraform experience for provisioning and managing AWS data infrastructure. •Strong Python skills for pipeline development, data transformation, and automation scripting. •Understanding of data reliability engineering, including idempotency, exactly-once processing, late-arriving data, schema evolution, and SLA-driven pipeline design. Preferred Qualifications •Experience in transportation, logistics, trucking, or fleet management, or with high-volume transactional SaaS platforms processing operational telemetry data. •Experience building DaaS platforms or data products, including governed external APIs, Redshift Data Sharing, or AWS Data Exchange integrations. •Knowledge of Parquet, ORC, and data lake or lakehouse patterns using Amazon S3 with Redshift Spectrum or AWS Glue. •Familiarity with BI and analytics tools such as Tableau, Power BI, Amazon QuickSight, or Looker, and an understanding of how data models affect query performance. •Experience with data observability and quality tools such as Monte Carlo, Great Expectations, or dbt tests. •Experience developing reusable data platform tooling, shared dbt packages, or internal data engineering frameworks. •Experience working with fully remote, distributed engineering teams.