Senior Solutions Architect – Large Scale AI Inference (España)

Senior Solutions Architect – Large Scale AI Inference (España)

09 sep
|
NVIDIA
|
España

09 sep

NVIDIA

España

We are looking for a Senior Solutions Architect with deep experience in large-scale production AI inference. You will collaborate with top EMEA AI Natives, AI infrastructure providers, and enterprises deploying AI at scale. As a trusted technical leader, you will ensure NVIDIA's inference stack achieves best performance, efficiency, and reliability. You will address the industry's toughest AI inference challenges at the intersection of AI and high-performance computing, driving innovations in areas such as multi-node Mixture-of-Experts (MoE) serving, interconnect-aware scheduling, memory-bound workload optimization, and next-generation inference architectures. Your work will establish the technical direction for scalable, high-performance AI inference across the most demanding production environments.

What You Will Be Doing

Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters.
Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs.




Improve inference efficiency across quantization (INT4/FP8), speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments.
Collaborate with NVIDIA product teams (Dynamo, TensorRT-LLM, NIXL) to accelerate customer success.
Animate the AI inference developer’s community across EMEA through technical workshops, hackathons, and reference architectures.

What We Need To See

MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience.
5+ years of experience in Neural Networks inference optimization.
Solid understanding of transformers inference optimization: quantization, disaggregated inference, speculative decoding, continuous batching, KV cache optimization.
Practical experience in MoE inference at scale: expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale.
Abi

📌 Senior Solutions Architect – Large Scale AI Inference (España)
🏢 NVIDIA
📍 España

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior solutions architect – large scale ai inference (españa) / españa

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior solutions architect – large scale ai inference (españa) / españa