04 oct
|
Lever
|
Cataluña
Finom’s AI Team seeks a seasoned data scientist to own evaluation methodologies for its internal AI agents. You will design benchmarks, set quality gates, and partner with domain experts to label cases and resolve disagreements.
You’ll work with Databricks, Claude Code, and dbt to build robust evaluation pipelines. You will lead the offline and online evaluation suites across ~20 processes and translate findings into actionable product decisions, balancing quality, cost, and latency.
#J-18808-Ljbffr
📌 Senior ai evaluation & quality scientist (eu remote) (Cataluña)
🏢 Lever
📍 Cataluña