Summary
A senior data engineering position responsible for designing and maintaining secure cloud-based data platforms. The role focuses on building scalable pipelines, improving data quality, supporting analytics and machine learning initiatives, and implementing modern data engineering practices.
Highlights
Remote senior data engineering opportunity focused on cloud architecture, analytics platforms, machine learning support, and building scalable data solutions for complex investigations and reporting needs.
Description
Hope you are doing well.
We have an open position for a Senior Data Engineer.
Pl.
go through the below description and let me know your interest.
If you are interested, kindly share a copy of your resume to hr@vinsysinfo.com along with the rate / salary and best time to reach you.
Role: Senior Data Engineer (Need 2 Candidates)
Work Arrangement: Remote/telework, with onsite participation when requested
Client: Federal Government SBA Office of Inspector General
Position Summary
The Senior Data Engineer will design, implement, maintain, and improve an integrated and flexible data architecture within SBA OIG s Microsoft Azure environment.
The role will support audits, investigations, fraud analytics, and machine-learning activities by developing sustainable data pipelines, migrating source data, improving data quality, implementing source control, and maintaining reliable cloud-based data-processing environments.
Responsibilities
Provide authoritative expertise in data-engineering methods and best practices.Apply code-first development approaches and modern pipeline-design patterns.Design and maintain a secure, stable, scalable, and flexible data architecture.Manage data assets through source control.Design, implement, and maintain ELT/ETL pipelines.Develop pipelines using Azure Synapse and Azure Machine Learning.Work with Azure Machine Learning SDK V1 and SDK V2.Migrate source data into Azure Data Lake Storage.Review, maintain, and improve existing architecture and pipelines.Conduct periodic reviews to identify bottlenecks, deprecated dependencies, and architecture drift.Implement pipeline quality controls, error handling, logging, monitoring, and validation checks.Incorporate source control into data pipelines and analytics codebases.Optimize data ingestion, processing, storage, and retrieval.Work with structured, semi-structured, and unstructured data.Use modern columnar formats, including Parquet.Normalize common entity attributes, including names, addresses, telephone numbers, and other identifying information.Develop self-service capabilities that allow SBA OIG analysts to query and export data.Coordinate with data scientists to support machine-learning models and analytical pipelines.Develop SOPs for authoring, developing, validating, publishing, executing, and monitoring pipelines and assets.Develop data dictionaries, entity-relationship diagrams, pipeline maps, and architecture documentation.Expand the environment with additional datasets and services as requested.Establish intake, testing, and production-deployment procedures.Monitor pipelines to ensure performance and regular dataset updates.Recommend architecture changes that reduce cloud costs.Evaluate emerging AI, automation, coding-assistant, and LLM-assisted data-engineering capabilities.
Required Qualifications
Candidates Must Possess One Of The Following
Bachelor s degree in Data Engineering, Computer Science, Data Science, Machine Learning, Mathematics, or a related field; orFive years of applied work experience in one or more of these fields.
Candidates must have at least five years of hands-on experience in each of the following:
Maintaining SQL databases.Conducting advanced SQL and T-SQL operations.Designing, implementing, and maintaining ELT/ETL processes in cloud-based data-analytics environments.
Candidates must have at least three years of hands-on experience in each of the following:
Working with Azure Synapse.Working with Azure Machine Learning.Working with modern data-stack technologies.Manipulating data using Python.Using Pandas.
Preferred Qualifications
Microsoft DP-203 certification or equivalent.PySpark or Polars experience.Experience developing reusable and modular code.Experience implementing pipelines and infrastructure using Python SDKs, command-line tools, REST APIs, or Infrastructure-as-Code tools.Experience implementing source-control and CI/CD workflows.Familiarity with AI coding assistants and LLM integration patterns.