13 sep
|
NVIDIA
|
San José de la Rinconada
13 sep
NVIDIA
San José de la Rinconada
NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the software powering today’s most sophisticated AI systems — from large language models to multimodal generative AI — all accelerated on NVIDIA GPUs. The Deep Learning Inference team develops and optimizes open-source frameworks that make AI deployment scalable, efficient, and accessible — including SGLang, vLLM, and FlashInfer. Our work enables developers worldwide to harness NVIDIA accelerators for real-time inference at every scale, from datacenter clusters to edge devices.
What you'll be doing:
Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering.
Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.
Oversee performance tuning, profiling,
and optimization of large-scale models for LLM, multimodal, and generative AI applications.
Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).
Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA’s broader AI and software strategies.
Foster a culture of technical excellence, open collaboration, and continuous innovation.
What we need to see:
MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
6+ overall years of software development experience, including 3+ years in technical leadership or engineering management.
Strong background in C/C++ software design and development; proficiency in Python is a plus.
Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.
Proven record of deploying or optimizing
📌 Engineering Manager, Deep Learning Inference (San José de la Rinconada)
🏢 NVIDIA
📍 San José de la Rinconada