30 sep
|
European Tech Recruit
|
España
30 sep
European Tech Recruit
España
Network Engineer - AI Infrastructure Fabric (IB/Spectrum-X, BGP)
Our client is a fast-scaling AI infrastructure operator building out next-generation GPU compute campuses across multiple European countries, supporting some of the largest AI training and inference workloads in the region. They're in the middle of an aggressive multi-site build-out, six deployments across four countries, and need experienced network engineers who can design and operate the high-speed fabric that ties tens of thousands of GPUs together.
The Role
You'll own the design, build, and operation of the network fabric connecting large-scale GPU clusters, including InfiniBand / NVIDIA Spectrum-X fabrics, plus the BGP routing at the network edge that connects these sites to the wider internet and to each other. This is hands-on, high-stakes infrastructure work: the fabric you build directly determines how efficiently thousands of GPUs can train and serve AI models.
What You'll Do
- Design, deploy, and operate InfiniBand and/or NVIDIA Spectrum-X fabric across multiple GPU cluster deployments
- Own BGP routing and edge network architecture across sites in Germany, the Netherlands, and Spain
- Work directly with platform and facility engineering teams during new site bring-up, ensuring fabric performance meets the demands of GB200/GB300-class GPU hardware
- Troubleshoot fabric-level performance issues affecting distributed training and inference workloads
- Build and maintain network automation to manage fabric configuration at scale
- Participate in an on-call rotation supporting live production fabric across sites
What We're Looking For
- Experience with InfiniBand and/or NVIDIA Spectrum-X fabric design and operation at scale
- Strong BGP and EVPN routing expertise, ideally including spine-leaf architectures
- Experience with fabric management tooling (e.g. NVIDIA UFM or equivalent)
- Network automation experience (Python, Ansible, or similar)
- Comfort working in a fast-moving, multi-site build-out environment with tight deadlines
- Prior experience at a hyperscaler, neocloud/GPU-cloud provider, or NVIDIA/Mellanox background is a strong plus.
Nice to Have
- Experience with NVIDIA DGX/HGX or GB200/GB300-class platform deployments
- Background supporting AI/ML training or inference infrastructure specifically
Why This Role
This is a rare opportunity to build fabric-level infrastructure for genuinely frontier-scale AI compute - tens of thousands of GPUs, exabyte-scale storage, and the newest hardware generation, during a period of hypergrowth, with real ownership over the network architecture from day one.
#J-18808-Ljbffr
📌 Network Engineer (España)
🏢 European Tech Recruit
📍 España