03 oct
|
Jobgether
|
España
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Engineer, LLM Inference Optimization based in Spain. As a Senior Machine Learning Engineer, you will drive the optimization of large language and vision-language model inference from model artifacts through production deployment. You will work across model internals, inference engines, serving architectures, and benchmarking to improve latency, throughput, memory efficiency, GPU utilization, reliability, and cost per token. This is a hands-on role focused on solving complex performance challenges and delivering measurable improvements to production systems. You will collaborate closely with kernel, platform, infrastructure, research, product, and customer-facing teams. Your work will involve evaluating serving configurations, diagnosing performance and quality regressions, and implementing advanced inference optimization techniques.
You will also establish reproducible benchmarks and safe rollout practices for high-throughput AI workloads. Accountabilities: - Own optimization initiatives for specific model families, customer endpoints, and inference serving backends. - Evaluate inference engines and recommend practical serving configurations based on workload requirements. - Diagnose and resolve model quality, performance, and reliability regressions during production rollouts. - Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token. - Deploy, configure, benchmark, and extend modern inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or equivalent technologies. - Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
- Implement or integrate advanced inference techniques such as speculative decoding,
📌 Senior Machine Learning Engineer, LLM Inference Optimization (España)
🏢 Jobgether
📍 España