AI Platform Operations Engineer - L2

Silicontek — United States · Posted ~3 hours ago

Mid Contract Hybrid

Skills

L2 operations production support Neo4j Palantir Foundry platform monitoring incident response troubleshooting cloud platforms AI platform operations AI platforms

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an L2 AI Platform Operations Engineer for a long-term hybrid role supporting production data and AI infrastructure. You will monitor platform health, troubleshoot failures, manage deployments and permissions, support incidents, and build dashboards, alerts, automated checks, and operational documentation.

Highlights

Long-term L2 operations role supporting production data, cloud, graph, and AI platforms. The position offers hands-on responsibility for platform reliability, monitoring, incident response, automation, dashboards, and operational best practices.

Description

Location: Fremont, CA (Only Local Candidates/Hybrid role) Duration Long Term Position profile: Provides consolidated L2 operations support for Neo4j, Palantir Foundry, and agents running on Foundry, covering platform availability, performance, deployments, permissions, and production stability. Key Responsibilities Monitor platform health, capacity, performance, scheduled jobs, pipelines, agents, and deployment status.Troubleshoot failures across data connections, builds, compute, permissions, and platform configuration.Distinguish platform, data, integration, and application-code issues and route them correctly.Support incident response, change control, post-incident reviews, and problem management. Build operational dashboards, alerts, automated health checks, and clear runbooks.Required skills and experience: Hands-on experience operating production data, cloud, graph, or AI platforms.Working knowledge of Neo4j and Palantir Foundry operations, including jobs, pipelines, compute, connectivity, permissions, and deployment support.Strong troubleshooting skills with Python and SQL for diagnostics and automation.Experience with monitoring, alerting, incident management, SLAs/SLOs, and on-call support.Ability to work across engineering, infrastructure, security, and vendors.