Summary
✨ AI‑Generated
A senior platform engineering opportunity focused on designing and maintaining scalable AWS infrastructure for data-intensive and AI-driven workloads. You will build automation, CI/CD foundations, internal tooling, and reliable production environments while partnering with engineering and data teams. The role suits an experienced engineer comfortable with cloud architecture, distributed systems, security, and operational excellence.
Highlights
Build and operate scalable AWS infrastructure supporting data platforms, AI workloads, automation, and production systems. The role offers broad ownership across infrastructure, data platforms, internal tooling, and reliability.
Description
Platform Engineer IV Location:- Denver, CO(Onsite) JOB SUMMARY Client’s Infrastructure Intelligence and Analytics (IIA) team builds and operates the data platform and AI agent infrastructure that powers proactive network monitoring and autonomous investigation for Client's network operations.
As part of this group, the Platform Engineer IV designs, builds, and maintains the AWS infrastructure that underpins the IIA Data Lake, agent runtime environments, CI/CD pipelines, and graph database systems, and develops the utilitarian application code, automation, and internal tooling for those systems.
This role ensures production environments are stable, scalable, and secure while enabling data science and agentic AI workloads to operate reliably at scale.
MAJOR DUTIES AND RESPONSIBILITIES
What you will do
span across the following areas.Individual focus areas will be determined based on team needs and candidate strengths: Infrastructure and Data Lake · Design and manage AWS infrastructure for the IIA Data Lake including S3 storage, Glue data catalog, Athena query engine, and EMR compute clusters.
· Manage cross-account connectivity, VPC networking, security groups, and IAM roles/policies to enable secure data flow between IIA, upstream data providers, and downstream consumers.
· Build and maintain infrastructure for AI agent runtime environments, including compute resources for LangGraph agents deployed via LangSmith Deployments.
· Support deployment and operation of AWS Neptune for the network topology graph (digital twin), including capacity planning, schema design support, and performance tuning.
· Implement and manage infrastructure-as-code (Terraform, CloudFormation) for repeatable, auditable environment provisioning.
· Manage IAM access key rotations, secrets management (AWS Secrets Manager, Delinea), and security compliance for on-premises and cloud integrations (e.g., Splunk Edge Processor).CI/CD and Agent Deployments · Build and maintain CI/CD pipelines for AI agent deployments using GitLab CI/CD, Docker, and Artifactory.
· Manage container lifecycle for agents deployed via LangSmith Deployments, including image builds, versioning, and rollback procedures.
· Automate deployment workflows to enable rapid, reliable promotion of agents from development through production.
· Coordinate with SpecGPT platform team on AI Gateway integration, cross-account deployment, and connectivity requirements.Application Development and Tooling · Develop and maintain utilitarian application code: scripts, CLIs, small services, and automation utilities (primarily Python) that support data ingestion, deployment, environment provisioning, and operational workflows.
· Write integration code and glue services that connect IIA systems with upstream data providers, downstream consumers, and external platforms.Production Operations · Ensure production environment stability through monitoring, alerting, and incident response.Maintain SLAs for data pipeline availability and agent uptime.
· Implement production monitoring and alerting for deployed agents (health checks, error rates, latency, resource utilization).
· Coordinate with upstream data teams and platform teams (SpecGPT, Splunk, Public Cloud) on connectivity, firewall requests, and integration requirements.
· Support data engineering team with infrastructure needs for new data source onboarding (storage provisioning, access controls, pipeline compute).
· Perform other duties as required.
REQUIRED QUALIFICATIONS Skills/Abilities and Knowledge · Ability to read, write, speak and understand English · Strong communication skills with ability to explain infrastructure decisions to non-infrastructure stakeholders · Expert-level experience with AWS services: EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatch · Strong experience with infrastructure-as-code (Terraform preferred, CloudFormation acceptable) · Experience managing cross-account AWS architectures, VPC peering, PrivateLink, and transit gateway configurations · Experience with IAM policy design, least-privilege access patterns, and service account management
Rajmurat Yadav E:rajmurat.y@t3pillars.com | 23710 Schooler Plaza Office 2024, Ashburn | VA - 20148