HPC Administrator
You will support the design, build, and operation of Linux-based high-performance computing clusters (CPU and GPU). You will manage system administration tasks, ensure cluster stability, maintain software stacks, and provide support for HPC users and scientific computing environments.
Key Responsibilities
- Design and build Linux-based HPC CPU/GPU clusters.
- Perform system administration, software maintenance, system monitoring, and troubleshooting for HPC/GPU clusters.
- Manage and operate HPC facilities and infrastructure.
- Administer Red Hat Enterprise Linux 8, 9, 10 or similar operating systems.
- Configure and support parallel computing environments CUDA, OpenMP, MPI.
- Deploy and manage Infiniband (ultra low latency networks) and Ethernet high-performance networks.
- Administer batch scheduling systems such as Slurm.
- Manage compute resources (CPU, GPU, RAM) for serial and parallel workloads.
- Compile, install, update, and tune scientific software and libraries (commercial and open‑source) on compute nodes and HPC file systems.
- Maintain and update firmware and drivers related to HPC systems.
- Provide user support for HPC environments, including troubleshooting, software assistance, guidance on cluster usage, and performance optimization.
Requirements
- Experience designing and managing HPC Linux clusters (CPU/GPU).
- Strong background in Linux system administration (preferably RHEL‑based).
- Knowledge of parallel programming environments (CUDA, OpenMP, MPI).
- Experience administering Slurm and managing HPC compute resources.
- Hands‑on experience with Infiniband and high‑performance networking.
- Proficiency with compiling and maintaining scientific software stacks.
- Strong troubleshooting skills in complex HPC environments.
Nice to Have
- Familiarity with HPC storage systems.
- Scripting knowledge (Bash, Python).
- Experience in performance tuning of HPC environments.
Location
Madrid
Work model
Hybrid Project
Why join us
Work at
📌 HPC Administrator (España)
🏢 Atos
📍 España