Senior Data Scientist (Ai Evaluation & Quality) (Barcelona)

Senior Data Scientist (Ai Evaluation & Quality) (Barcelona)

08 oct
|
Finom
|
Barcelona

08 oct

Finom

Barcelona

You'll join Finom's AI Team as the founding IC dedicated to the quality, evaluation and telemetry of AI agents powering Finom's internal operations and tools (including Ops workflows and internal AI analytics engines across ~20 core processes) Own and extend our offline/online evaluation suites across ~20 internal AI agent processes-datasets (capability + regression), LLM-as-a-judge rubrics, and deterministic checks Establish pre-launch quality gates: enforce pass/fail thresholds in CI/CD pipelines before agent prompt, context, or tool changes hit production Work directly with domain experts to label cases and resolve annotator disagreement-breaking definition criteria rather than averaging disagreement away Build test datasets derived from real user & operational traffic (tickets, internal chats, colleague queries) rather than synthetic edge cases Harden statistical methodology: handle judge drift, verbosity bias, non-determinism, and measure true metric shifts vs. noise Translate quality numbers into operational decisions: run weekly syncs with process owners to define clear quality vs. cost/latency trade-offs Statistical Rigor: Applied knowledge of sampling,



hypothesis testing, variance analysis, and confidence intervals on noisy metrics Production LLM Experience: In the last 1-2 years, you have built, shipped, or evaluated LLM-based systems (RAG, multi-step tool use, agents) as a core, primary job responsibility5+ years in Data Science / Product Analytics / Applied AI roles, with sustained product-level metric ownership Fluent Python & SQL: Ability to write clean data pipelines, evaluation harnesses, and dbt transformation models directly Autonomous Quality Ownership: Proven track record of owning evaluation methodology or analytics for an entire product or end-to-end process (what to build vs. what NOT to build)AI-assisted coding (Claude Code, Cursor, or Codex) is your default daily authoring environment for Python, SQL, and evaluation scripts-not something you occasionally experiment with You can walk us through concrete work tasks from the last month where AI coding tools accelerated your engineering and data analysis #J-18808-Ljbffr

📌 Senior Data Scientist (Ai Evaluation & Quality) (Barcelona)
🏢 Finom
📍 Barcelona

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior data scientist (ai evaluation & quality) (barcelona) / barcelona

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior data scientist (ai evaluation & quality) (barcelona) / barcelona