Experteer Overview In this role you will own and advance the workload orchestration stack across HPC and AI platforms, enabling scalable, efficient execution of massive scientific and AI workloads. You’ll work within the Accelerated Compute Engineering team to optimize resource use and ensure high availability in multi-tenant environments. You’ll bridge traditional HPC scheduling with modern AI orchestrators to accelerate research and development at Roche. This opportunity offers hands-on impact in a mission-driven, collaborative setting that values innovation and reliability.Compensaciones / Ventajas
- Design, implement, and maintain SLURM Workload Manager ecosystem across HPC clusters for high availability and efficient resource distribution
- Deploy and manage Run:ai as core orchestration and virtualization layer for AI Factory with fractional GPU allocation
- Evaluate and implement SLURM Slinky integrations to connect Kubernetes-based AI orchestration with HPC resources
- Define best practices for containerized execution using Singularity/Apptainer and Enroot for reproducible HPC environments
- Tune scheduling parameters (topology-aware scheduling, multi-node scaling), QoS,
and fair-share policies to maximize multi-tenant efficiency
- Collaborate on observability with monitoring and telemetry dashboards to track scheduler efficiency and hardware utilization
- Troubleshoot complex workload failures (distributed training, MPI bottlenecks, driver issues)
- Maintain configuration-as-code models for scheduling policies and automated deployment of cluster configurationsResponsabilidades
- 5+ years of systems engineering focusing on workload scheduling, resource management, and multi-tenant cluster optimization
- Expert-level proficiency with SLURM (partition design, accounting, plug-ins)
- Strong experience with Singularity for container execution
- Hands-on experience with Run:ai, Kubernetes, and containerized GPU scheduling
- Solid understanding of InfiniBand/ RoCE and multi-node communication (MPI, NCCL)
- Automation proficiency in scheduler configurations and telemetry tooling
- Bachelor's or higher in Computer Science, Applied Mathematics, Computational Engineering, or similarRequisitos principales
-
📌 Workload Orchestration Engineer (Madrid)
🏢 Roche
📍 Madrid
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.