Summary
✨ AI‑Generated
A senior engineering role responsible for operating highly available cloud-based data and AI platforms. The position focuses on infrastructure automation, reliability engineering, observability, security, and supporting generative AI workloads.
Highlights
Remote senior engineering opportunity working on modern cloud and AI platforms. The role offers ownership of reliability, automation, security, and performance improvements for advanced data environments.
Description
Senior/ Lead Data Engineer- Snowflake Cortex+ AWS Bedrock
Years of Experience: 6- 12 years
Location: Remote
W2 or Independent Contactors
We are looking for an experienced Site Reliability Engineer (SRE) to support and operate highly available Snowflake and AWS-based data and AI platforms, with hands-on exposure to Snowflake Cortex and Amazon Bedrock.
The engineer will be responsible for platform reliability, monitoring, incident management, automation, performance optimization, security, and operational support for data and Generative AI workloads.
Key Responsibilities
Design, deploy, monitor, and maintain highly available AWS, Snowflake, and GenAI platforms.Support Snowflake Cortex capabilities for AI/ML and Generative AI workloads.Work with Amazon Bedrock to integrate and operate foundation models and GenAI applications.Monitor infrastructure, applications, data pipelines, APIs, and AI workloads.Define and monitor SLIs, SLOs, and SLAs for critical services.Implement observability using CloudWatch, logs, metrics, traces, and alerting.Troubleshoot production incidents and perform Root Cause Analysis (RCA).Participate in on-call / production support and incident management.Automate repetitive operational activities using Python, Shell scripting, Terraform, or AWS services.Build and maintain CI/CD pipelines for application, data, and AI workloads.Optimize Snowflake compute usage, query performance, warehouse utilization, and cost.Monitor and troubleshoot Snowflake data pipelines and workloads.Support AWS services such as S3, Lambda, IAM, CloudWatch, ECS/EKS, API Gateway, and Secrets Manager.Implement security, access controls, encryption, secrets management, and least-privilege IAM.Monitor the reliability and performance of LLM/GenAI applications using Amazon Bedrock.Track AI application metrics such as latency, throughput, errors, token usage, and cost.Establish automated health checks, alerts, recovery mechanisms, and disaster-recovery procedures.Collaborate with Data Engineers, ML Engineers, Developers, Architects, Product Managers, and business stakeholders.Continuously improve platform reliability through automation, observability, capacity planning, and performance engineering.