Senior Cloud & Machine Learning Engineer

Testingxperts — United States · Posted ~1 day ago

Senior Full-time Onsite

Skills

AWS Python Amazon SageMaker CloudFormation Terraform Infrastructure as Code DevOps Machine Learning Engineering Cloud Architecture Jenkins GitHub Actions

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior engineering opportunity for an expert in cloud infrastructure and machine learning workflows. You will architect and automate scalable environments, develop Python-based model pipelines, manage infrastructure as code, and collaborate with data and ML specialists to move models from development into reliable production operations.

Highlights

Senior technical role combining cloud architecture, machine learning engineering, and DevOps automation. Offers hands-on work with advanced cloud and AI/ML services, infrastructure as code, and production-oriented model deployment workflows.

Description

Houston, TX Role Summary Seeking a senior-level technical resource with deep AWS expertise (Level 300+) to support cloud infrastructure development, machine learning engineering workflows, and DevOps automation. This resource will work closely with Data Scientists and ML Engineers to accelerate model development, deployment, and operationalization on AWS. Required Capabilities AWS Platform Expertise (L300+) Deep, hands-on knowledge of core and advanced AWS services including compute, storage, networking, security, and AI/ML services. Ability to architect solutions, interpret service quotas/limits, and apply AWS Well-Architected best practices. SageMaker Development (Python) Proficient in writing Python code for SageMaker workflows - training jobs, processing jobs, model hosting, endpoint deployment, hyperparameter tuning, pipelines (SageMaker Pipelines), and experiment tracking. Infrastructure as Code (IaC) Ability to author and maintain CloudFormation templates and Terraform modules for provisioning AWS resources (VPCs, IAM roles, S3 buckets, SageMaker domains, Lambda functions, Step Functions, etc.) with parameterization and environment-based configuration. Programmatic AWS Resource Deployment (Python) Experience writing Python scripts (Boto3/AWS SDK) for automated resource provisioning, configuration, and lifecycle management - custom deployment scripts, one-off migrations, and resource orchestration. Advanced Troubleshooting (L300+) Ability to diagnose and resolve complex issues across AWS services and local/remote development environments - including IAM permission errors, networking misconfigurations, service integration failures, SDK/CLI issues, and containerized workload debugging. ML & AI Engineering Support Hands-on support for Data Scientists and ML Engineers in model building, fine-tuning, evaluation, and deployment. Experience with agent-based architectures - including Bedrock Agents, AgentCore runtime (managed agent hosting, MCP server integration, session management), LangChain, and custom orchestration frameworks - and MLOps patterns. CI/CD Automation Working knowledge of CI/CD pipelines using Jenkins and GitHub Actions - including build/test/deploy stages, artifact management, environment promotion, and integration with AWS services (CodePipeline, ECR, ECS/EKS deployments). Ideal Background Must have hands-on AWS experience across multiple service domains Strong Python development skills with production-quality coding practices Familiarity with ML lifecycle tooling (SageMaker, MLflow, or equivalent) Experience with agentic AI frameworks and managed agent runtimes (AgentCore, Bedrock Agents) Experience embedding within cross-functional teams (Data Science, Platform Engineering) Certifications Required: AWS Solutions Architect Associate, AWS Certified Generative AI Developer Preferred: Solutions Architect Professional, ML Specialty, or DevOps Engineer