Systems Engineer - Cloud Operations

Autozone — United States · Posted ~1 day ago

Skills

cloud infrastructure Google Cloud Platform Terraform Kubernetes GitOps CI/CD observability cloud operations infrastructure automation reliability engineering cloud security GKE Argo CD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A systems engineering role focused on deploying, managing, and optimizing cloud infrastructure. You will automate infrastructure provisioning, operate Kubernetes-based environments, manage GitOps and CI/CD workflows, implement observability, and support infrastructure for AI/ML applications. Close collaboration with development, reliability, and platform engineering teams is central to the role.

Highlights

Cloud operations role focused on building reliable and secure infrastructure, automation, Kubernetes, GitOps, observability, and emerging AI/ML platform initiatives.

Description

As a Systems Engineer on the Cloud Operations team, you will be responsible for deploying, managing, and optimizing our cloud-based infrastructure on Google Cloud Platform (GCP). You will work with technologies such as Terraform, Kubernetes (GKE), GitOps/ArgoCD, CI/CD pipelines, and observability tools to ensure reliable, secure, and scalable platform operations. You will also contribute to our AI/ML platform initiatives, supporting infrastructure for LLM-based applications and AI-powered automation tools that enhance developer productivity and operational efficiency. You will collaborate with development teams, SREs, and platform architects to ensure seamless deployment and delivery of applications while maintaining the highest standards of reliability, security, and performance. Cloud Infrastructure, Automation & Operations Design, build, and maintain cloud infrastructure using Terraform to automate provisioning, scaling, and lifecycle management of resources on GCPDevelop and maintain CI/CD pipelines using GitLab CI to automate build, test, and deployment workflows. Implement and maintain GitOps practices using ArgoCD for declarative, version-controlled application deploymentMonitor system performance using observability tools (Dynatrace, Cloud Monitoring, Prometheus/Grafana) and troubleshoot production issuesParticipate in on-call rotation to provide 24/7 support for critical infrastructure incidentsPerform root cause analysis on incidents and implement preventive measures. Document runbooks, architecture decisions, and operational procedures Kubernetes Platform Management Deploy, configure, and manage containerized applications on Google Kubernetes Engine (GKE), including GKE Autopilot and Standard clusters Manage cluster lifecycle including upgrades, node pool configurations, and capacity planningTroubleshoot pod failures, CrashLoopBackOff, OOMKilled events, and container resource issuesConfigure and optimize resource requests/limits, Horizontal Pod Autoscaler (HPA), and Vertical Pod Autoscaler (VPA)Manage Kubernetes networking including Services, Ingress controllers, Network Policies, and DNS configurations. Implement and manage service mesh (Istio) for traffic management, observability, and securityManage secrets and configurations using Kubernetes Secrets, ConfigMaps, and external secret management tools. Implement pod security standards, RBAC policies, and workload identity configurations AI/ML Platform & Automation Support infrastructure for AI/ML workloads including LLM-based applications and model serving platformsDeploy and manage AI-powered developer tools such as coding assistants (Claude Code, GitHub Copilot) and agentic AI systems. Explore and implement AI-assisted incident response and automated remediation workflowsBuild and maintain infrastructure for Retrieval-Augmented Generation (RAG) pipelines and vector databasesConfigure GPU-enabled node pools and optimize resource allocation for AI/ML workloadsImplement MCP (Model Context Protocol) servers and AI agent integrations for operational automationStay current with emerging AI technologies and evaluate their applicability for infrastructure automation Kubernetes Expertise (Essential) 3+ years hands-on experience with Kubernetes in production environmentsDeep understanding of Kubernetes architecture: API server, etcd, scheduler, controller manager, kubeletExperience with GKE (Standard and Autopilot modes), including cluster creation, upgrades, and maintenanceProficiency in troubleshooting workloads: analyzing pod logs, events, describe outputs, and container statesStrong understanding of resource management: requests, limits, QoS classes, and resource quotasExperience with Kubernetes networking: Services (ClusterIP, NodePort, LoadBalancer), Ingress, Network PoliciesKnowledge of Kubernetes storage: PersistentVolumes, PersistentVolumeClaims, StorageClasses, dynamic provisioningExperience with Helm charts for application packaging and deploymentFamiliarity with Kubernetes security: RBAC, Pod Security Standards, Secrets management, Workload IdentityUnderstanding of Kubernetes observability: metrics-server, kubectl top, container resource monitoringExperience debugging common issues: ImagePullBackOff, CrashLoopBackOff, OOMKilled, Evicted pods, pending pods Cloud & Infrastructure 3+ years of experience with Google Cloud Platform (GCP) services including GKE, Cloud Run, Cloud SQL, Memorystore, Pub/Sub, and Cloud LoggingStrong experience with Terraform for infrastructure as code (IaC)Understanding of cloud networking: VPCs, subnets, firewall rules, Cloud NAT, Private Service Connect CI/CD & GitOps Proficiency with GitLab CI/CD pipelinesExperience with ArgoCD or similar GitOps toolsUnderstanding of Helm charts and Kustomize for Kubernetes manifest management Observability & Troubleshooting Experience with monitoring and APM tools (Dynatrace, Datadog, Prometheus, Grafana)Ability to analyze logs, metrics, and traces to diagnose production issuesFamiliarity with JVM troubleshooting (heap dumps, thread analysis, GC tuning, connection pool issues) AI/ML Knowledge Basic understanding of LLM concepts, prompt engineering, and AI model deploymentFamiliarity with AI coding assistants and their integration into development workflowsInterest in agentic AI systems and autonomous automation toolsExposure to vector databases (Pinecone, Weaviate, pgvector) and RAG architectures is a plus Systems & Networking Strong Linux administration skillsUnderstanding of networking concepts (DNS, load balancing, firewalls, TCP/IP)Experience with service mesh (Istio) is a plus General Excellent problem-solving and analytical skillsStrong written and verbal communicationAbility to work effectively in a collaborative, cross-functional environmentExperience working in an Agile/DevOps cultureBachelor's degree in Computer Science, Information Technology, or related field (or equivalent experience)