05 sep
|
Capitole
|
España
DATA SCIENTIST – DOCUMENT INTELLIGENCE
About the role
We are looking for a Data Scientist to join an international project focused on developing and improving AI-powered document processing solutions.
In this role, you will work with LLMs, Vision-Language Models, OCR, NLP and Machine Learning to transform complex unstructured business documents into structured data.
You will collaborate with AI Engineers, Data Engineers, Product teams and business stakeholders to develop and evaluate solutions that deliver measurable business value.
If you enjoy solving complex Data Science challenges and experimenting with AI to solve real business problems, this could be your next challenge!
What you'll do
Design and improve AI solutions for document extraction, classification, validation and transformation.
Experiment with LLMs, multimodal models, OCR, NLP and Machine Learning to identify the best approach for each use case.
Develop prompts, structured outputs, validation and post-processing logic for complex document formats.
Build datasets and evaluation frameworks to measure accuracy and continuously improve solution performance.
Develop automated evaluations and regression tests to ensure robustness when models or pipelines change.
Write clean, reusable Python code and translate successful experiments into production-ready solutions with AI and Data Engineers.
Work with Product and business stakeholders to define use cases, success criteria and measurable AI solutions.
Must Have
3+ years of experience in Data Science, Machine Learning or Applied AI.
Hands-on experience with LLMs and Generative AI , including prompt engineering and experimentation.
Strong understanding of NLP, Machine Learning and statistics , including when classical or deterministic approaches are more appropriate than LLMs.
Experience working with unstructured or semi-structured data , datasets, ground truth and evaluation metrics.
Strong Python skills and good practices around testing, version control and reproducible experimentation.
Experience analysing results and translating findings into measurable improvements.
Fluent English and spanish and ability to communicate technical concepts to both technical and non-technical stakeholders.
Degree in Computer Science, AI, Data Science, Mathematics, Engineering or a related field.
✨ Nice to Have
Experience with Document Intelligence and OCR , particularly Azure Document Intelligence .
Experience with managed AI platforms such as Azure OpenAI or Amazon Bedrock .
Experience with document classification, information extraction or NER .
Experience applying multimodal or Vision-Language Models to document content and layouts.
Hybrid model – 2 days on site per week
Why join this project?
📌 Data Scientist – Document Intelligence (España)
🏢 Capitole
📍 España