Responsibilities
Own and extend our offline eval suite across products — datasets (capability + regression), judges, metrics
Build and maintain online quality dashboards: resolution rate, CSAT, thumbs up/down, LLM-as-judge signals, error rate, latency
Close the production feedback loop: mine failure patterns from real traffic → turn them into regression cases → propose fixes to Product and domain experts
Harden methodology: judge stability, non-determinism handling
Translate numbers into decisions – weekly syncs, clear trade‑offs, no dashboards for their own sake
Qualifications
Solid foundation in statistics — sampling, hypothesis testing, variance, understanding what a noisy metric is
Python and SQL — you can build an analysis end-to-end
3+ years in analyst / data scientist roles, at least one in a product context
Analytical mindset — you start from the business question, not from the tool
Experience in quality analytics for ML systems — ranking, recommendations, classification, etc.
Hands‑on experience evaluating LLM applications (RAG, agents, tool use, judges)
Experience building LLM agents — side projects, toy builds, personal experiments all count
📌 Product Data Scientist (AI Evaluation & Quality) (Madrid)
🏢 Finom
📍 Madrid