Senior AI Platform Engineer

Shinebask Technologies Llc — United States · Posted ~1 day ago

Senior Full-time Hybrid

Skills

AI platform engineering AI agents CI/CD Infrastructure as Code Cloud infrastructure Secure deployment Identity and access management Secrets management Observability Incident response Automation Technical leadership AI Cloud

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking a Senior AI Platform Engineer to implement its responsible AI and agentic systems strategy. You will build scalable platform capabilities, automate provisioning and CI/CD, establish secure runtime architectures, improve observability, and provide technical leadership across engineering teams in a hybrid full-time role.

Highlights

Lead hands-on engineering for scalable and responsible AI infrastructure, with benefits and a full-time role. Influence engineering standards, improve deployment automation, and help teams operate secure, reliable AI systems.

Description

Full Time Role: Senior AI Platform Engineer _ Richardson , TX Location: Hybrid (Richardson , TX) Duration FTE + Benefits Overview: This Senior Engineer drives the implementation of an organization's responsible AI and agentic systems strategy through scalable infrastructure, automated delivery pipelines, production-grade operational practices, and engineering standards. The role combines hands-on AI platform engineering with governance, observability, secure deployment, and technical leadership across engineering teams. Responsibilities Core Responsibilities Define standards-based deployment patterns for AI agents, reusable platform capabilities, and secure runtime architectures, including identity, access, and secrets management.Automate AI platform provisioning, standardized release processes, environment management, CI/CD pipelines, infrastructure-as-code deployments, and operational documentation.Operate and maintain reliable AI platform services through incident response, capacity planning, monitoring, observability, rollback procedures, and disaster recovery capabilities.Collaborate with security, cloud infrastructure, DevOps, and platform engineering teams to establish governance, guardrails, secure release practices, observability standards, and production readiness requirements.Mentor engineers, lead technical design reviews, guide implementations toward established architectural standards, and balance tradeoffs among cost, velocity, reliability, scalability, and security. Qualifications Required Experience 8-10 years of software engineering experience, including at least 7 years building and operating production systems in cloud environments and hands-on experience deploying AI/ML services into production.2-3 years of experience utilizing AI-assisted development tools (for example, Claude Code, Codex, Cursor, GitHub Copilot, or similar) to accelerate platform development, deployment, testing, troubleshooting, and documentation.Strong experience with AWS, GCP, Azure, hybrid cloud, or on-premises infrastructure; Docker, Kubernetes, GitHub Actions or equivalent CI/CD tooling, Terraform, and Helm.Working knowledge of authentication and authorization technologies, including OAuth 2.0, OpenID Connect (OIDC), SAML, JWT, Role-Based Access Control (RBAC), and Identity and Access Management (IAM).Experience implementing workload identity, service-to-service authentication, and securing API and tool access.4+ years of scripting and automation experience, preferably with Python and JavaScript, including troubleshooting across Linux, containers, Kubernetes, networking, and distributed production systems.Ability to architect secure and cost-efficient hosting solutions for open-source large language models (LLMs) in cloud, private cloud, or on-premises environments. Preferred Qualifications Experience with agentic AI frameworks (e.g., LangGraph, Google ADK), agent-to-agent communication patterns, tool integration, workflow orchestration, Retrieval-Augmented Generation (RAG), AI evaluation frameworks, and guardrail design.Familiarity with AI observability and tracing platforms (e.g., LangSmith, Grafana/LGTM), GPU infrastructure for model serving, and orchestration platforms such as Dagster, Prefect, or Apache Airflow.Strong communication, mentoring, stakeholder management, and project planning skills.Ability to support critical production deployments and incident response activities when necessary. Success Metrics Reusable deployment patterns are documented, standardized, automated, observable, and auditable across AI and agent-based workloads.Development, testing, and production release processes become more efficient, secure, consistent, and easier to govern.Security, compliance, governance, observability, and responsible automation practices are embedded throughout the software delivery lifecycle. Compensation & Benefits The organization offers a competitive compensation package and comprehensive benefits, which may include: Medical, dental, and vision insuranceRetirement savings programs with employer contributionsPaid time off and company holidaysProfessional development and continuing education opportunitiesPerformance-based incentive or bonus programs, where applicable