06 oct
|
Lever
|
Barcelona
About FinomFinom is a European tech startup headquartered in Amsterdam, and we're on a journey towards revolutionizing the financial landscape for entrepreneurs worldwide. Our mission is to develop an all-in-one financial B2B solution that integrates banking functions, accounting, financial management, and invoicing into a seamless, mobile-first platform. This significant investment follows a $105 million growth funding round from General Catalyst, a long-term backer since 2021 known for supporting companies like Airbnb, HubSpot, KAYAK, and Stripe.
Finom's platform goes beyond traditional banking, offering invoicing and a growing suite of features, including AI-enabled accounting, aiming to simplify financial management for entrepreneurs. We nurture innovation and an inspiring work environment where bold ideas thrive, prioritizing thorough research, swift implementation of solutions, and ensuring that every effort we make benefits our users, employees, partners, and our business as a whole. Maintaining our start-up spirit, we prioritize thorough research, swift implementation of solutions, and ensuring that every effort we make benefits our users, employees, partners, and, of course, our business.
AI TeamYou'll join Finom's AI Team as the founding IC dedicated to the quality, evaluation and telemetry of AI agents powering Finom's internal operations and tools (including Ops workflows and internal AI analytics engines across ~20 core processes).Our beliefAn AI agent is only as good as the evaluation loop running on it. Your missionDesign the evaluation methodology, build golden benchmarks,
and establish quality gates for our internal AI agents from scratch—working directly with process owners and domain experts, with no senior quality owner above you to lean on. Core StackDatabricks, DeepEval, Claude Code, Cursor, Python, SQL, dbt.
What You Will Be
DoingOwn and extend our offline/online evaluation suites across ~20 internal AI agent processes—datasets (capability + regression), LLM-as-a-judge rubrics, and deterministic checks. Establish pre-launch quality gates: enforce pass/fail thresholds in CI/CD pipelines before agent prompt, context, or tool changes hit production. Build test datasets derived from real user & operational traffic (tickets, internal chats, colleague queries) rather than synthetic edge cases.
Harden statistical methodology: handle judge drift, verbosity bias, non-determinism, and measure true metric shifts vs. noise. Translate quality numbers into operational decisions: run weekly syncs with process owners to define clear quality vs. Must-Haves5+ years in Data Science / Product Analytics / Applied AI roles, with sustained product-level metric ownership.
Autonomous Quality Ownership: Fluent Python & SQL: Ability to write clean data pipelines, evaluation harnesses, and dbt transformation models directly.
Statistical Rigor: Applied knowledge of sampling, hypothesis testing, variance analysis,
and confidence intervals on noisy metrics.
Daily
Setup & ToolingAI-assisted coding (Claude Code, Cursor, or Codex) is your default daily authoring environment for Python, SQL, and evaluation scripts—not something you occasionally experiment with. You can walk us through concrete work tasks from the last month where AI coding tools accelerated your engineering and data analysis. How we work — one thing we mean seriouslyAI-assisted coding is our default authoring environment, not a bonusClaude Code is our main tool — you'll reach for it for SQL, Python, analyses, dashboards, and internal scriptsWe're looking for analysts who are already curious and fluent with AI coding — or genuinely excited to become fluent fastWe care about what you ship and how clearly you thinkIf this idea excites you rather than worries you, you'll feel at home hereWhat You Will Get In ReturnMake a genuine impact on the product Join our upward trajectory, and grow with us.
We provide the resources and opportunities for continuous personal and professional development, empowering you to make a genuine impact on our evolving product. Work in the EU Embark on this exciting journey with us and enjoy the flexibility of traveling and working remotely or in a hybrid model across Europe. This exciting opportunity is available to every team member, from junior team members to our founders.
It's the idóneo opportunity to strike the perfect work-life balance while enjoying breathtaking Mediterranean views.
Equal Opportunity StatementAt Finom, we're an equal opportunity employer and value diversity at our company.
📌 Senior Data Scientist — AI Evaluation & Quality | Internal AI Agents (Remote in Europe) (Barcelona)
🏢 Lever
📍 Barcelona