Senior IT Infrastructure Specialist

Cohesity — United States · Posted ~3 hours ago

Senior Full-time

Skills

IT infrastructure Cloud platforms System administration Cybersecurity Cloud Infrastructure Automation

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior infrastructure engineering role responsible for supporting secure, scalable environments across production and development platforms. The position requires strong technical troubleshooting and operational skills.

Highlights

Work on modern infrastructure environments with exposure to cloud, security, automation, and emerging technologies.

Description

Cohesity is a leader in AI-powered data security and management. Aided by an extensive ecosystem of partners, Cohesity makes it easy to secure, protect, manage, and get value from data — across the data center, edge, and cloud. Cohesity helps organizations defend against cybersecurity threats with comprehensive data security and management capabilities, including immutable backup snapshots, AI-based threat detection, monitoring for malicious behavior, and rapid recovery at scale. We’ve been named a Leader by multiple analyst firms and have been globally recognized for Innovation, Product Strength, and Simplicity in Design. Join us on our mission to shape the future of our industry. Senior IT Infrastructure Engineer About The Role We are seeking a Senior IT Infrastructure engineer who thrives in dynamic environments and can rapidly learn new platforms, embrace AI-assisted operations, and contribute across both lab and production infrastructure. Success in this role requires strong technical depth, operational ownership, collaboration, and a bias toward automation and continuous improvement. You'll work across Linux systems, Kubernetes clusters, kubevirt-based virtualization, and enterprise SAN storage. You will be expected to use AI tools to move faster, reduce toil, and catch issues earlier than a purely manual workflow would allow. This is not an "AI-optional" role. We expect modern infrastructure engineers to pair deep systems knowledge with AI-assisted workflows for scripting, troubleshooting, documentation, and operational analysis. The judgment and systems expertise are still yours; AI is a force multiplier, not a replacement for understanding what's happening under the hood. What You'll Do Design, deploy, and maintain Linux server infrastructure (RHEL/Rocky/Ubuntu and similar) across production and non-production environmentsOperate and scale Kubernetes clusters: deployments, networking, storage integration, upgrades, and troubleshootingManage kubevirt-based virtualization environments: VM lifecycle, resource allocation, live migration, performance tuning, and host maintenanceAdminister and optimize SAN storage systems (Pure Storage, Dell, HPE, and/or NetApp), including provisioning, performance, snapshots, and capacity planningContribute to datacenter networking design and troubleshooting (VLANs, switching, routing fundamentals) where applicableBuild automation and tooling to reduce repetitive manual work across the infrastructure stackParticipate in on-call rotation and incident response, driving root-cause analysis for infrastructure issuesDocument architecture, runbooks, and operational proceduresParticipate in maintenance activities, platform upgrades, and infrastructure migrations across global data center locations How AI Fits Into This Role We expect you to actively use AI tools to augment (not replace) the skills above. This looks like: Scripting & automation: Using AI coding assistants (e.g., Claude Code, Copilot) to draft and refine automation for provisioning, patching, and cluster operations; reviewing and validating output rather than blindly trusting itKubernetes operations: Using AI to help generate and audit services and troubleshoot pod/network/storage issues faster by summarizing logs and correlating events across clustersIncident response: Using AI to accelerate log analysis, correlate symptoms across systems (compute, storage, network) during outages, and draft initial incident timelines/postmortemsCapacity & performance analysis: Using AI-assisted data analysis to identify SAN storage trends, VM resource pressure, or cluster bottlenecks before they become incidentsDocumentation: Using AI to turn tribal knowledge into clear runbooks, architecture diagrams, and onboarding docs, keeping documentation current with less manual overheadKnowledge gaps: Using AI as a first-pass research tool for unfamiliar storage arrays, network gear, or Kubernetes add-ons, then validating against vendor docs and testing in a safe environment We care about outcomes and judgment, not tool usage for its own sake. You should be comfortable explaining why an AI-suggested command, config, or fix is correct before running it in production. Requirements 5+ years of hands-on experience in Linux systems administrationStrong production experience with Kubernetes (deployment, operations, troubleshooting at scale)Solid experience with kubevirt-based virtualizationComfortable using AI tools (coding assistants, LLM-based troubleshooting/research) as part of daily engineering workProven experience operating production infrastructure supporting critical internal business services with defined uptime, performance, and recovery objectives. Preferred SAN storage management experience with Pure Storage, Dell, HPE, and/or NetApp platformsDatacenter networking experience (switching, routing, VLANs)Experience building internal tooling or automation that incorporates AI/LLM APIs What Success Looks Like Expected to become an active contributor to production infrastructure operations within the first 60-90 days, including participation in incident response, operational reviews, and infrastructure change managementOperate across multiple technology domains.Incidents are resolved faster because AI-assisted analysis shortens the diagnostic loopManual, repetitive operational work steadily decreases as automation (AI-assisted or otherwise) takes it overThe team's institutional knowledge is captured in living documentation rather than a single person's head Disclosure Pursuant to Applicable State Equal Pay Transparency Laws - This position has a starting pay range as listed below. Actual salary depends upon many factors, including a candidate’s skills, qualifications and experience, location, and salary expectations, and therefore a starting salary at the low end, high end, or even above the stated range may be offered. This position may also be eligible for bonus compensation, commission (if in a sales function), and/or equity grants. Additionally, full-time employees are eligible to participate in our comprehensive benefits framework, including health and wellness benefits, vacation, paid holidays and refresh days, 401(k) retirement plan, life and disability insurance coverages, and other benefits the Company may offer from time to time. Pay Range : $101,320.00-$126,650.00 The compensation noted above is based on an annualized hourly rate assuming normal full-time employment. Data Privacy Notice for Job Candidates: For information on personal data processing, please see our Privacy Policy. Equal Employment Opportunity Employer (EEOE) Cohesity is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status or any other category protected by law. If you are an individual with a disability and require a reasonable accommodation to complete any part of the application process, or are limited in the ability or unable to access or use this online application process and need an alternative method for applying, you may contact us at 1-855-9COHESITY or recruiting@cohesity.com for assistance. In-Office Expectations Cohesity employees who are within a reasonable commute (e.g. within a forty-five (45) minute average travel time) work out of our core offices 2-3 days a week of their choosing. We strongly prefer candidates who are currently located in or near the designated job location. Candidates outside the area should apply only if they are committed to relocating prior to their start date and have the legal right to work in the job location.