Summary
✨ AI‑Generated
A DevOps engineering position focused on scalable infrastructure for AI-driven applications, automation, reliability improvements, and collaboration with engineering and security teams.
Highlights
Build reliable infrastructure for AI workloads and improve software delivery through automation and cloud-native engineering.
Description
Job Description:
We are seeking a highly skilled DevOps / Site Reliability Engineer (SRE) with experience supporting modern AI platforms and cloud-native infrastructure.
This role will focus on building scalable, reliable infrastructure for AI workloads while partnering closely with security and engineering teams to operationalize findings from an emerging AI-driven security platform used to identify code vulnerabilities and infrastructure risks.
This position is ideal for an engineer who enjoys automating infrastructure, improving software delivery pipelines, and supporting the rapid adoption of AI technologies in enterprise environments.
Responsibilities:Design, build, and maintain highly available infrastructure supporting AI and machine learning platforms.Develop scalable platform engineering solutions that enable reliable deployment and operation of AI services.Partner with development and security teams to remediate vulnerabilities and infrastructure issues identified by the AI driven security platform.Improve platform reliability through automation, monitoring, observability, and proactive performance tuning.Build and maintain robust CI/CD pipelines for application and infrastructure deployments.Automate operational workflows using Python and Infrastructure-as-Code practices.Implement DevSecOps best practices throughout the software development lifecycle.Support containerized workloads and cloud-native applications.Troubleshoot production issues, perform root cause analysis, and implement long-term reliability improvements.Optimize deployment strategies, release automation, and infrastructure scalability.Collaborate with AI engineering teams to ensure AI services are secure, resilient, and production-ready.Required Qualifications:5+ years of experience in DevOps, Site Reliability Engineering, or Platform EngineeringStrong experience designing and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or similar)Strong Python scripting and automation skillsExperience supporting cloud infrastructure (AWS, Azure, or GCP)Experience with Infrastructure as Code (Terraform, CloudFormation, or Pulumi)Hands-on experience with Docker and KubernetesStrong understanding of Linux systems administrationExperience implementing monitoring and observability solutions (Prometheus, Grafana, Datadog, Splunk, etc.)Experience working with security scanning tools and vulnerability remediationFamiliarity with DevSecOps principles and secure software deliveryExperience supporting AI platform engineering or machine learning infrastructurePreferred Qualifications:Understanding of AI model deployment, inference infrastructure, and scalability considerationsFamiliarity with GPU-enabled infrastructure and AI compute environmentsKnowledge of vector databases, LLM deployment, or MLOps conceptsExperience with Kubernetes operators, service mesh, or distributed systemsExposure to AI security tooling Experience integrating automated security scanning into CI/CD pipelinesRole requires 3 days onsite per week
US persons only
Rate: $65/hr.
W2 or $75/hr.
Corp.