Ml Engineer For Llm Inference (Madrid)

Ml Engineer For Llm Inference (Madrid)

03 oct
|
HireHi
|
Madrid

03 oct

HireHi

Madrid

ОписаниеSocial Discovery Group is a group of social discovery companies that addresses loneliness, isolation, and disconnection by transforming virtual intimacy into the new normal. Its products are social entertainment platforms designed to connect people online across different cultures and regions.ЗадачиSpeed up and scale LLM inference in production using SGLang, KV and prefix caching, batching, quantization, and speculative decodingRun distributed inference for very large models with up to 1T+ parameters across multi-GPU and multi-node setupsBenchmark new GPU servers and hardware, bring them into production, and adapt serving code to themLead the NLP and CV teams technically by reviewing experiments, setting direction, and stepping in early when neededTrain and fine-tune language models, and improve the agent harnesses and chat algorithm that run on themTrack cutting-edge research and open-source work in inference and post-training, and turn it into the ML roadmapCollaborate with validation, content, and dataset preparation teams to design experiments and measure model qualityТребованияDeep hands-on experience optimizing LLM inference in production with SGLang, vLLM, or TensorRT-LLMExperience with distributed inference or training of large models, including MoE, tensor/expert/pipeline parallelism, and multi-node GPU clustersStrong understanding of inference performance, including KV cache, attention kernels, batching, quantization, and GPU profilingExperience training and fine-tuning LLMs, including post-training such as RLHF or DPOProven technical leadership, including guiding engineers through reviews, mentoring,



and technical decisions while continuing to write codeProficiency with PyTorch, transformers, and related librariesAdvanced English or RussianБудет плюсом: experience at AI-focused startups or companies such as Character AI or OpenAI, backend engineering experience with Python, Go, or C#, knowledge of scalable deployment systems, CUDA or Triton kernel development, a computer vision background, experience accelerating large generative image or video models, experience with multimodal LLMs, first-author papers or notable open-source work such as contributions to SGLang, vLLM, or post-training libraries, a degree in CS, math, or physics from a strong program (MSc or PhD)УсловияThe initial pay level or pay range for this role will be shared with candidates during the recruitment process and before the commencement of employment28 Calendar days of vacation per year7 Wellness days per year that can be used to deal with household issues or recover without taking sick leaveBonuses up to $5000 for recommending successful applicants for positions in the company50% Payment for professional training, international conferences, and meetingsCorporate discount for English lessonsHealth benefits: if not eligible for corporate medical insurance, employees receive compensation of up to $1,000 gross per year for health insurance or doctors’ fees for themselves and close relativesThe company provides equipped workplaces and necessary equipment in its offices or co-working locations; elsewhere, it reimburses workplace costs up to $1,000 gross once every 3 years#J-18808-Ljbffr

📌 Ml Engineer For Llm Inference (Madrid)
🏢 HireHi
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ml engineer for llm inference (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ml engineer for llm inference (madrid) / madrid