Roche seeks a seasoned systems engineer to lead SLURM-based workload scheduling across multi-tenant HPC clusters. You will implement Run:ai for fractional GPU allocation and bridge Kubernetes-based orchestration with traditional HPC resources.
You will optimize queues, QoS, and fair-share policies while ensuring high availability and reproducible environments using Singularity/Apptainer. Idóneo candidates will have deep Linux, SLURM expertise, and hands-on experience with MPI/NCCL for distributed
📌 AI & HPC Workload Orchestrator (Madrid)
🏢 Roche
📍 Madrid