16 sep
|
Next-link
|
Barcelona
16 sep
Next-link
Barcelona
At NextLink, we are looking for a Data Scientist – Generative AI to join our team for a 6–12 month collaboration, working closely with one of our key clients in the pharmaceutical sector.
You will join the iQAgent and Unstructured Data Pipeline area, focusing on improving the quality of an AI-based search and question-answering solution for enterprise documents.
The role combines hands-on Data Science, applied Generative AI engineering, RAG and retrieval optimization, evaluation design, and technical guidance, with the goal of improving answer quality and generating evidence to support strategic decisions around the future search architecture.
Main Responsibilities:
- Assess the existing iQAgent and document-processing pipeline, focusing on answer quality, prompt behavior, retrieval performance, agent workflows, and document extraction quality.
- Define and operationalize evaluation metrics, datasets, baselines, and repeatable testing procedures covering retrieval quality, faithfulness, grounding, source quality, and answer quality.
- Implement and validate improvements to prompts, retrieval logic, ranking, context construction, and agent workflows in collaboration with engineering teams.
- Apply feature extraction, metadata enrichment, and signal generation to unstructured documents to improve search, retrieval, ranking, and answer generation.
- Review text and image extraction quality within the document pipeline and recommend practical improvements.
- Use evaluation evidence to prioritize improvements based on expected user value and provide pragmatic technical guidance to developers.
- Document findings and communicate clear, actionable recommendations to Product, Architecture, and Engineering stakeholders.
Requirements
- Strong Python development and data analysis skills, with hands-on experience building or improving applied AI solutions.
- Strong knowledge of LLMs, RAG, and agentic AI, including practical experience with frameworks such as LangChain, LangGraph, or similar.
- Comprehensive understanding of LLM and RAG evaluation, including metrics, test design, evaluation datasets, baselines, and reproducible measurement approaches.
- Practical experience with evaluation and observability frameworks such as MLflow, Langfuse, DeepEval, OpenAI Evals, or comparable solutions.
- Solid understanding of information retrieval, vector search, embeddings, document chunking, ranking, and context construction.
- Hands-on prompt engineering experience, including optimization of system prompts, user prompts, tool instructions, and context for enterprise AI agents.
- Experience with feature extraction, metadata enrichment, and signal generation from unstructured documents to improve retrieval and answer quality.
- Ability to review, prototype, and implement improvements within an existing technical solution, while providing pragmatic guidance to engineering teams.
- Strong analytical mindset and ability to use objective evidence and evaluation results to prioritize quality improvements.
- Excellent communication and presentation skills in English, both written and spoken, with the ability to communicate effectively with Product, Architecture, and Engineering stakeholders.
📌 Data Scientist – Generative AI (Barcelona)
🏢 Next-link
📍 Barcelona