Principal Cloud Engineering and Production Operations Engineer

A10Networks — United States · Posted ~2 hours ago

Lead Full-time

Skills

Cloud Infrastructure DevOps Security Multi-Cloud AWS Azure Kubernetes Terraform

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Seeking a Principal Cloud and Production Operations Engineer to architect and optimize hybrid cloud environments for customer-facing applications. The role involves designing scalable, secure, and resilient systems while fostering a culture of automation and observability. Collaborate with DevOps, Security, and IT teams to ensure enterprise-grade infrastructure reliability.

Highlights

Lead cloud-native infrastructure design, ensure system scalability and security, and drive automation initiatives in a high-tech IT environment.

Description

The Principal Cloud and Production Operations Engineer serves as the senior technical authority responsible for architecting, automating, and optimizing hybrid and cloud-native production environments that power critical customer-facing services and enterprise applications. This role combines deep cloud infrastructure expertise with strong production reliability and operational engineering skills. The Principal Engineer acts as both architect and hands-on builder, ensuring scalability, resilience, and security across multi-cloud and on-prem environments. Reporting to the Associate Director of IT and Infrastructure, this position will collaborate closely with Engineering, DevOps, Security, and IT Operations to drive a culture of automation, observability, and continuous improvement across the production ecosystem. Key Responsibilities Cloud Architecture and Engineering Design, implement, and maintain cloud and hybrid infrastructure supporting production workloads, enterprise systems, and CI/CD pipelinesLead the adoption of infrastructure-as-code (IaC) using Terraform, CloudFormation, or similar tools to enable repeatable, auditable, and secure deploymentsArchitect scalable and fault-tolerant solutions across OCI, AWS, Azure, and on-prem data centers, ensuring high availability and cost efficiencyEvaluate emerging cloud services and technologies for applicability to business needs and long-term scalability goals Production Operations and Reliability Serve as the technical lead for production operations, ensuring uptime, performance, and reliability of customer-facing and internal systemsDevelop and maintain observability frameworks leveraging metrics, logs, and traces to ensure proactive detection and rapid responsePartner with engineering teams to implement SRE-inspired practices, including service level objectives (SLOs), error budgets, and post-incident reviewsDrive root cause analysis, performance tuning, and continuous improvement of production services Automation and CI/CD Enablement Collaborate with DevOps and application engineering teams to build and optimize automated deployment pipelines supporting frequent, low-risk releasesIntegrate security and compliance checks into CI/CD workflows to ensure production readiness and alignment with internal standardsDesign self-healing infrastructure and automated rollback mechanisms to reduce operational riskEnsure secure and reliable configuration management and environment orchestration using tools such as Ansible, Chef, or Puppet Operational Governance and Collaboration Establish and enforce operational best practices for monitoring, patching, and change management across production systemsLead production readiness reviews for new releases and large-scale changesCollaborate with the Security and Compliance teams to ensure systems adhere to policy, hardening standards, and regulatory requirementsParticipate in and occasionally lead on-call rotations for critical production systems, ensuring rapid triage and resolution Leadership and Mentorship Act as a technical mentor to cloud and infrastructure engineers, fostering a culture of knowledge sharing and engineering excellenceLead architectural reviews, design sessions, and capacity planning discussionsServe as a trusted advisor to management on cloud modernization, resilience engineering, and cost optimization strategies Qualifications Bachelor’s degree in Computer Science, Information Systems, or related field; Master’s preferred10+ years of experience in cloud and infrastructure engineering, including 3+ years in a senior or principal roleExpertise with OCI (preferred), AWS and/or Azure cloud services, including networking, compute, storage, and identity managementProven experience managing production-scale environments supporting mission-critical applications and servicesStrong proficiency in:Infrastructure-as-code (Terraform, CloudFormation)CI/CD and DevOps toolchains (Jenkins, GitLab, ArgoCD)Container orchestration (Kubernetes, Docker)Monitoring and observability platforms (Prometheus, Grafana, Datadog, ELK)Scripting and automation (Python, Bash, PowerShell)Solid understanding of security, compliance, and networking principles in hybrid environmentsExceptional analytical, problem-solving, and incident management skillsDemonstrated ability to lead complex, cross-functional initiatives from concept to execution Preferred Experience Experience in high-availability SaaS or networking environmentsKnowledge of FinOps, cost optimization, and multi-cloud governance frameworksFamiliarity with Zero Trust, identity federation, and cloud access security modelExposure to AI/ML infrastructure or data-driven pipelines is a plus AI Use Guidelines for Interviews: Our interviews are designed to reflect your own skills and thinking. The use of AI or recording tools during live interviews is not permitted unless explicitly invited by the interviewer or approved in advance as part of a reasonable accommodation. If these tools are used inappropriately or in a way that misrepresents your work, your application may not move forward in the process. Why Join Us This is a hands-on leadership opportunity to define the next generation of cloud and production operations within a high-impact technology environment. The Principal Cloud and Production Operations Engineer will directly influence the reliability, speed, and scalability of the company’s global technology platforms ensuring operational excellence and innovation. A10 Networks is an equal opportunity employer and a VEVRAA federal subcontractor. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability status, protected veteran status, or any other characteristic protected by law. A10 also complies with all applicable state and local laws governing nondiscrimination in employment. Hybrid Targeted compensation guideline: $140,000 - $185,000. Compensation will vary based on number of factors, including market demand for specific skills, role type, job level, and individual qualifications. Final salary offers are determined by considerations including, but not limited to, subject matter expertise, demonstrated skill level, relevant experience, geographic location, education, certifications, and training.