Senior MLOps Engineer (Training & Inference Optimization) (San Sebastián)

Senior MLOps Engineer (Training & Inference Optimization) (San Sebastián)

04 ago
|
Multiversecomputing
|
San Sebastián

04 ago

Multiversecomputing

San Sebastián

ph3Multiverse Computing /h3 pMultiverse is a well-funded, fast-growing deep-tech company founded in 2019. We are the largest quantum software company in the EU and have been recognized by CB Insights (2023 and 2025) as one of the 100 most promising AI companies in the world. With 180+ employees and growing, our team is fully multicultural and international. We deliver hyper-efficient software for companies seeking a competitive edge through quantum computing and artificial intelligence. Our flagship products, CompactifAI and Singularity, address critical needs across various industries: /p ul liCompactifAI is a groundbreaking compression tool for foundational AI models based on Tensor Networks. It enables the compression of large AI systems—such as language models—to make them significantly more efficient and portable. /li liSingularity is a quantum- and quantum-inspired optimization platform used by blue-chip companies to solve complex problems in finance, energy, manufacturing, and beyond. It integrates seamlessly with existing systems and delivers immediate performance gains on classical and quantum hardware. /li /ul pYou’ll be working alongside world‑leading experts to develop solutions that tackle real‑world challenges. We’re looking for passionate individuals eager to grow in an ethics‑driven environment that values sustainability and diversity. We’re committed to building a truly inclusive culture—come and join us. /p h3About the Role /h3 pWe are seeking a bSenior MLOps Engineer /b to steer the technical vision of our Training and Inference Optimization team. In this high‑impact role, you will architect the infrastructure that powers our next‑generation AI models. You will bridge the gap between systems programming and machine learning, optimizing large‑scale LLM training via bNVIDIA NeMo /b and building ultra‑high‑throughput serving systems using bvLLM /b, bTensorRT-LLM /b, and bSGLang /b. Your mission is to ensure our models are not only state‑of‑the‑art but also production‑hardened, cost‑efficient, and performant at scale. /p h3Key Responsibilities /h3 ul libTraining Infrastructure:



/b Architect and maintain scalable distributed training pipelines using bNVIDIA NeMo/Nemotron/Megatron‑Bridge /b. You will optimise GPU utilisation, manage complex checkpointing strategies, and implement automated fault tolerance for long‑running jobs. /li libInference Orchestration: /b Lead the deployment of LLMs using bvLLM, TensorRT‑LLM, or SGLang /b. You will implement and tune cutting‑edge techniques—including bPagedAttention /b, continuous batching, and advanced quantisation (bAWQ/FP8 /b)—to maximise throughput and minimise bTPOT /b (Time Per Output Token). /li libWorkload Orchestration: /b Utilize bSLURM/Flyte/Ray/SkyPilot /b to manage and scale ML workloads across diverse cloud providers and on‑prem clusters, ensuring seamless resource shifting and cost‑effective execution. /li libLifecycle Management: /b Standardise model tracking, versioning, and transition workflows using bMLflow /b (or similar tool), ensuring reproducible training runs and a clear path from research to production. /li libPerformance Engineering: /b Conduct deep‑dive profiling and bottleneck analysis across the full stack—from bCUDA kernels /b and bNCCL /b collective communications to Python‑level orchestration. /li libEfficiency Cost Governance: /b Monitor and optimise cloud and on‑prem GPU expenditures through intelligent scaling policies and high‑density resource packing. /li libTechnical Leadership: /b Set the bar for engineering excellence. You will drive the roadmap, perform rigorous code reviews, and mentor junior and mid‑level engineers. /li /ul h3Required Qualifications /h3 ul libExperience: /b 5+ years in MLOps, DevOps, or Software Engineering, with a minimum of 2 years dedicated to bLLM infrastructure /b. /li libDeep Learning Ecosystem: /b Expert‑level proficiency with bPyTorch /b and the NVIDIA stack (bCUDA, NCCL, Triton /b).



/li libSpecialised Tooling: /b Hands‑on experience with bNVIDIA NeMo /b (or Megatron‑Bridge) for distributed training and at least two of the following for serving: bvLLM, TensorRT‑LLM, or SGLang /b. /li libOrchestration Lifecycle: /b Proven experience with bSLURM/Flyte/Ray/SkyPilot /b for cluster management and bMLflow /b (or similar tool) for experiment and model management. /li libInfrastructure: /b Deep expertise in bKubernetes /b and K8s operators (e.g., bKubeRay /b, bMPI Operator /b, or bRun:ai /b). /li libSystems Programming: /b Mastery of Python and a functional understanding of bC++ or Rust /b for performance‑critical components. /li libNext‑Gen Hardware: /b Familiarity with high‑performance networking (bInfiniBand/RoCE /b) and NVIDIA bH200/B200 (Blackwell) /b architectures. /li /ul h3Preferred Skills /h3 ul liActive contributions to relevant open‑source projects (bvLLM, SGLang, SkyPilot, or NeMo /b). /li liProven track record with model compression (Sparsity, Distillation, or Quantisation). /li liExperience writing or optimising custom bTriton kernels /b. /li /ul pExpertise in ML observability stacks (Prometheus, Grafana, Jaeger). /p h3Perks Benefits /h3 ul liIndefinite contract /li liEqual pay guaranteed. /li liVariable performance bonus. /li liSigning bonus. /li liWe offer work visa sponsorship (If applicable). /li liRelocation package (if applicable). /li liPrivate health insurance. /li liEligibility for educational budget according to internal policy. /li liHybrid opportunity. /li liFlexible working hours. /li liLanguage classes and discounted lunch options. /li liWorking in a high paced environment, working on cutting edge technologies. /li liCareer plan. Opportunity to learn and teach. /li /ul pAs an equal opportunity employer, Multiverse Computing is committed to building an inclusive workplace. The company welcomes people from all different backgrounds, including age, citizenship, ethnic and racial origins, gender identities, individuals with disabilities, marital status, religions and ideologies, and sexual orientations to apply. /p /p #J-18808-Ljbffr

📌 Senior MLOps Engineer (Training & Inference Optimization) (San Sebastián)
🏢 Multiversecomputing
📍 San Sebastián

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior mlops engineer (training & inference optimization) (san sebastián) / san sebastián

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior mlops engineer (training & inference optimization) (san sebastián) / san sebastián