ph3Responsibilities /h3 ul liOwn and extend our offline eval suite across products — datasets (capability + regression), judges, metrics /li liBuild and maintain online quality dashboards: resolution rate, CSAT, thumbs up/down, LLM-as-judge signals, error rate, latency /li liClose the production feedback loop: mine failure patterns from real traffic → turn them into regression cases → propose fixes to Product and domain experts /li liHarden methodology: judge stability, non-determinism handling /li liTranslate numbers into decisions – weekly syncs, clear trade‑offs, no dashboards for their own sake /li /ul h3Qualifications /h3 ul liSolid foundation in statistics — sampling, hypothesis testing, variance,
understanding what a noisy metric is /li liPython and SQL — you can build an analysis end-to-end /li li3+ years in analyst / data scientist roles, at least one in a product context /li liAnalytical mindset — you start from the business question, not from the tool /li liExperience in quality analytics for ML systems — ranking, recommendations, classification, etc. /li liHands‑on experience evaluating LLM applications (RAG, agents, tool use, judges) /li liExperience building LLM agents — side projects, toy builds, personal experiments all count /li /ul /p #J-18808-Ljbffr
📌 Product Data Scientist (AI Evaluation & Quality) (España)
🏢 Finom
📍 España