SRE Platform Engineer - AI Agents

Topsysitsolutions — United States · Posted ~3 hours ago

Senior Onsite

Skills

Site Reliability Engineering Platform engineering Cloud infrastructure Kubernetes Automation Monitoring SLA/SLO management Cloud platforms

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced platform engineering role focused on building reliable AI-enabled infrastructure. The position involves Kubernetes, cloud environments, automation, monitoring, and improving production reliability.

Highlights

Advanced infrastructure role combining reliability engineering, cloud platforms, automation, and emerging AI technologies.

Description

Job Title: SRE Platform Engineer – AI Agents Location: Atlanta, GA (Inperson Interview) Job Summary We are looking for an experienced SRE Platform Engineer with 8+ years of experience in Site Reliability Engineering, platform engineering, cloud infrastructure, Kubernetes, and automation. The ideal candidate will also have hands-on exposure to AI Agents / Agentic AI and be comfortable working on modern AI-driven platform solutions. The candidate should have strong knowledge of SRE principles, SLA/SLO, Kubernetes, cloud platforms, automation, monitoring, and production support, along with solid programming and problem-solving skills. Key Responsibilities Design, build, maintain, and improve highly reliable and scalable platform infrastructure.Implement Site Reliability Engineering (SRE) practices across production environments.Define and monitor SLAs, SLOs, SLIs, error budgets, and service reliability metrics.Develop and maintain Kubernetes-based applications and platform infrastructure.Troubleshoot production issues, perform root-cause analysis, and implement long-term corrective actions.Build automation for deployment, monitoring, infrastructure management, and operational processes.Work with AI Agents / Agentic AI solutions and integrate AI capabilities into platform and operational workflows.Support CI/CD pipelines and DevOps automation.Monitor application and infrastructure health using logging, metrics, and alerting tools.Collaborate with development, DevOps, cloud, and AI engineering teams.Participate in technical design discussions and contribute to platform architecture.Write clean, efficient code/scripts and demonstrate strong problem-solving abilities.Required Skills 8+ years of experience in SRE, Platform Engineering, DevOps, or related infrastructure roles.Strong understanding of SRE concepts and practices.Strong knowledge of SLA, SLO, SLI, and Error Budgets.Hands-on experience with Kubernetes and containerized environments.Strong understanding of cloud infrastructure and production systems.Experience with CI/CD, automation, monitoring, logging, and incident management.Strong programming/scripting skills in languages such as Python, Java, Go, or similar.Experience with AI Agents / Agentic AI or AI-driven automation.Strong troubleshooting, debugging, and analytical skills.Excellent communication and collaboration skills.