05 ago
|
Talent
|
Valencia
pstrongOur client /strongis a fast-growing deep-tech company founded in 2019 and recognized by CB Insights as one of the 100 most promising AI companies globally. They are the largest quantum software company in the EU, with 250+ employees worldwide and growing, delivering advanced solutions trusted by leading general enterprises across several critical industries, including finance, energy, manufacturing, telecom, and industrial sectors. /ppbr/ppbr/ppstrongJob Overview /strong /ppThey are seeking a strongMLOps Engineer /strong to steer the technical vision of their Training and Inference Optimization team. In this high-impact role, you will architect the infrastructure that powers our next-generation AI models. You will bridge the gap between systems programming and machine learning, optimizing large-scale LLM training via strongNVIDIA NeMo /strong and building ultra-high-throughput serving systems using strongvLLM /strong, strongTensorRT-LLM /strong, and strongSGLang /strong. /ppYour mission is to ensure our models are not only state-of-the-art but also production-hardened, cost-efficient, and performant at scale. /p pbr/ppbr/ppstrongPerks and Benefits: /strong /pulliIndefinite contract. /liliEqual pay guaranteed. /liliVariable performance bonus. /liliSigning bonus. /liliThey offer work visa sponsorship (If applicable) and relocation package (if applicable). /liliPrivate health insurance. /liliEligibility for educational budget according to internal policy. /liliHybrid opportunity. /liliFlexible working hours. /liliLanguage classes and discounted lunch options. /liliA high-performance, collaborative environment, operating at pace on cutting-edge technologies. /liliCareer plan. Opportunity to learn and teach.
/li /ulpbr/ppbr/ppstrongRequired Qualifications: /strong /pullistrongExperience: /strong 5+ years in MLOps, DevOps, or Software Engineering, with a minimum of 2 years dedicated to strongLLM infrastructure /strong. /lilistrongDeep Learning Ecosystem: /strong Expert-level proficiency with strongPyTorch /strong and the NVIDIA stack (strongCUDA, NCCL, Triton /strong). /lilistrongSpecialized Tooling: /strong Hands-on experience with strongNVIDIA NeMo /strong (or Megatron-Bridge) for distributed training and at least two of the following for serving: strongvLLM, TensorRT-LLM, or SGLang /strong. /lilistrongOrchestration Lifecycle: /strong Proven experience with strongSLURM/Flyte/Ray/SkyPilot /strong for cluster management and strongMLflow /strong(or similar tool) for experiment and model management. /lilistrongInfrastructure: /strong Deep expertise in strongKubernetes /strong and K8s operators (e.g., KubeRay, MPI Operator, or Run:ai). /lilistrongSystems Programming: /strong Mastery of Python and a functional understanding of strongC++ or Rust /strong for performance-critical components. /lilistrongNext-Gen Hardware: /strong Familiarity with high-performance networking (strongInfiniBand/RoCE /strong) and NVIDIA strongH200/B200 (Blackwell) /strong architectures. /li /ulpbr/ppbr/ppstrongPreferred Qualifications: /strong /pulliActive contributions to relevant open-source projects (strongvLLM, SGLang, SkyPilot, or NeMo /strong). /liliProven track record with model compression (Sparsity, Distillation, or Quantization). /liliExperience writing or optimizing custom strongTriton kernels /strong. /liliExpertise in ML observability stacks (Prometheus, Grafana, Jaeger). /li /ul
📌 MLOps Engineer (LLM/NVIDIA) (Valencia)
🏢 Talent
📍 Valencia