Senior DevOps Engineer (HPC)

Epam Systems โ€” Poland ยท Posted ~2 days ago

Senior Hybrid Visa History โœ“

Skills

Terraform Linux HPC SRE Slurm Grafana Prometheus Ansible Bash Python OpenStack Salt Puppet

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Help build and operate large-scale research computing infrastructure using infrastructure-as-code, automation, cloud platforms, and reliability engineering while collaborating with technical and scientific teams.

Highlights

Work on advanced high-performance computing infrastructure, automation, and cloud technologies with opportunities for remote work, learning, certifications, and relocation.

Description

We are seeking a Senior DevOps Engineer to enhance our high-performance computing services and collaborate closely with the scientific community to optimize research computing. Join our team to build and operate cutting-edge HPC capabilities using automation and infrastructure-as-code. Apply now to contribute to innovative computational solutions in a dynamic environment. Responsibilities Design, implement, and maintain robust platform infrastructure using Infrastructure as Code tools such as TerraformDevelop, deliver, and operate research computing services and applicationsApply Site Reliability Engineering principles to manage HPC service deployment, monitoring, and incident responseSolve complex technical problems related to HPC services and user applicationsManage large-scale HPC, HTC, or BC computing environments for optimal performanceCollaborate with scientific users to tailor HPC resources to research needsAutomate deployment processes to ensure consistency across HPC infrastructureMaintain and administer large-scale cluster and server computing software such as Slurm, LSF, or Grid EngineDevelop and maintain monitoring dashboards using tools like Grafana and PrometheusWork within a DevOps team environment following agile methodologiesOperate and utilize virtualized private cloud resources such as OpenStackAdminister large-scale parallel filesystems including Weka, GPFS, or LustreUse configuration management tools like Ansible, Salt, or Puppet to manage IT operationsDevelop scripts and tools for HPC and DevOps platform operations using Bash and Python Requirements 3+ years of experience with DevOps processes and automation using Infrastructure as Code tools such as TerraformHands-on experience operating or engineering large-scale HPC or similar computing environmentsProven expertise in Linux system administration including TCP/IP networking and storage subsystemsExperience administering large-scale cluster management software such as Slurm, LSF, or Grid EngineKnowledge of configuration management tools like Ansible, Salt, or PuppetExperience working in agile DevOps teamsAbility to develop and maintain monitoring tools such as Grafana and PrometheusExperience with scripting languages such as Bash and Python for automation and tool developmentStrong experience managing virtualized private cloud environments like OpenStackScientific degree or equivalent experience in computationally intensive scientific data analysisProven ability to manage relationships with third-party suppliersUpper-intermediate proficiency in English (B2+) Nice to have Experience with container technologies such as LXD, Singularity, Docker, or KubernetesOperation and configuration experience with public cloud platforms like AWS, Azure, or GCPExperience with HashiCorp tools such as Vault, Consul, and NomadDevelopment experience with programming languages such as Java, C++, Python, Ruby, or PerlExperience with parallel filesystems like Weka, GPFS, or Lustre We offer We gather like-minded people:Top tech minds driving innovation in AI, cloud and digital platform modernizationSupportive team and agile, startup-like cultureHybrid by design mode and opportunity to work remotely within PolandChance to work abroad for up to 60 days annuallyBusiness-driven relocation opportunitiesWe provide growth opportunities:Career development programsThought leadership, mentoring, soft skills and well-being programsCertification (Anthropic, Gemini, GCP, Azure, AWS)English classesWe cover it all:Stable payParticipation in the Employee Stock Purchase Plan with a 15% discountBenefits package (health insurance, multisport, shopping vouchers)Referral bonuses up to $2,000Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and moreCorporate, social and well-being eventsPlease, note:Benefits listed above are available to employees onlyWe are open for working with Contractors. Terms of B2B cooperation agreements are agreed individuallyWe will reach out to selected candidates exclusively EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.