Lead Back End Engineer (Fully Remote) (Madrid)

Lead Back End Engineer (Fully Remote) (Madrid)

22 sep
|
Jobgether
|
Madrid

22 sep

Jobgether

Madrid

Our partner is looking for a Technical Lead - GPU Infrastructure based in Spain. You will lead the evolution from managed Kubernetes workloads toward bare-metal GPU infrastructure, including Slurm-based research computing and Kubernetes-powered inference. You will oversee a distributed team spanning backend, frontend, DevOps, QA, and documentation while maintaining high technical standards.

Your work will support research, model training, and managed inference workloads requiring reliable, scalable, and observable GPU compute. This is an opportunity to shape the architecture and operational foundations of a sophisticated GPU platform in a fully remote environment.

The Technical

Lead will own the platform architecture and engineering delivery while remaining deeply involved in technical decisions, infrastructure operations, team leadership, and partner relationships. Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation. Establish engineering standards, oversee code and design reviews, manage release gates, conduct one-to-ones, and provide growth and performance feedback.

Design, build, and operate a managed Slurm service supporting research and model-training workloads. Own Slurm controllers, accounting, partitions, login nodes, node onboarding, acceptance testing, driver and CUDA baselines, upgrades, stalled-job detection, node health, draining, autohealing, storage visibility, identity, and workload isolation. Lead GPU infrastructure operations on bare-metal environments, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.

Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator. Act as the primary technical interface with infrastructure partners and vendors,



translating requirements into written specifications and acceptance tests. Manage technical escalations with partners through resolution and contribute to capacity planning and hardware sourcing decisions.

Work directly with research, model-training, and product teams to translate workloads into platform requirements and manage capacity constraints. Contribute to architecture decisions involving distributed systems, high-performance computing, networking, storage, virtualization, and GPU workloads. Bachelor's or Master's degree in computer science, engineering, or a related field, or equivalent practical experience.

Experience operating HPC or GPU training clusters for research or model-development users. Strong experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycles, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance. Deep knowledge of InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.

Strong

Linux systems expertise, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning. Proven production Kubernetes experience covering control planes, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy.

Experience with HPC storage and large-scale data movement, including shared filesystems such as VAST, Lustre, or NFS and node-local NVMe caching.



Strong observability and operations experience with Prometheus, Grafana, Loki, or comparable platforms, including SLOs, incident response, and post-incident reviews. Working proficiency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions.

Excellent written and spoken English. Fully remote availability with a working location between UTC and UTC+5:30 to provide overlap with teams and partners across Europe and India. Slurm operators on Kubernetes, such as Soperator or Slinky, or Kubernetes-native schedulers such as Kueue, Volcano, KAI, or Kubeflow Trainer.

GPU parallelism strategies, quantization trade-offs, and GPU memory planning.

Experience working on the operator side of a GPU cloud, university or national HPC center, or AI research platform.

Experience working with hardware providers responsible for provisioning but not operating infrastructure, including establishing contracts and acceptance tests. 100% remote position. High-impact technical leadership role spanning bare-metal GPU infrastructure, Slurm, Kubernetes, inference, and observability. Significant ownership over architecture, engineering standards, delivery planning, and team development.

Opportunity to support advanced AI research, model training, and managed inference workloads. Opportunity to work at the intersection of high-performance computing, AI infrastructure, distributed systems, and cloud-native technologies.

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR).

📌 Lead Back End Engineer (Fully Remote) (Madrid)
🏢 Jobgether
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: lead back end engineer (fully remote) (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: lead back end engineer (fully remote) (madrid) / madrid