25 sep
|
haddock
|
Barcelona
- We’re looking for an AI Engineer who is analytical, quantitative and deeply AI-native. Someone who parallelizes hypotheses with coding agents, stays relentlessly focused on AI that scales and solves real customer pain, and measures before optimizing without ever letting analysis slow down shipping
- Document understanding at scale. Every invoice and delivery note a restaurant receives goes through our OCR. Accuracy at single field level, business casuistics for document understanding, cost per document and latency are all on the table, every single day
- Fina, our conversational agent, and the most ambitious thing we build. Fina sits on top of everything haddock knows about a restaurant: purchases, suppliers, stock, sales, bank movements, P&L.; Ask it anything about the business and it answers, and increasingly it goes and does the work instead of just reporting on it. It’s the surface where the whole system has to behave as one, and most of what it will become is still unbuilt
- End-to-end workflows that get the job done. Complex reconciliations, data workflows, and imports of the most chaotic data you can imagine: a supplier catalogue inside a photographed PDF, a POS export nobody ever documented, ten years of a restaurant’s history in a spreadsheet built by six different people. Somebody has to turn all of that into something a system can trust
- Product and format matching. Turning the messy free text of thousands of suppliers into a clean catalogue a restaurant can run its inventory on
- The internal packages every one of our agents shares: AI observability, evals and human annotation. They are ours, we maintain them, and the quality of every agent we ship depends on them
- Architect and ship AI solutions for restaurateurs, balancing latency, cost, and accuracy
- Obsess over efficiency. A single line of code has saved us thousands of dollars, minutes of latency, or unlocked a step change in accuracy. You’ll think this way by default, and you’ll keep the cost of what we run in view as we scale
- Own AI observability end to end: monitor billions of traces, design evaluation frameworks, and close the improvement loop as fast as possible
- Maintain and develop our internal packages for AI evals, AI observability, human annotation, domain . Every agent we run is built on top of them, so making them better makes everything better at once
- Build the harness. A model on its own is not a product. The tools it can call, the data it can read, the guardrails, the place where a human steps in: you build the backend, the frontend and the plumbing that turns a model into something thousands of restaurants rely on every day
- Research and experiment. New models and techniques land every few weeks and some of them change what’s possible for us. Running the experiment, reading the numbers and deciding whether it’s worth it is part of the job, not something you get to when there’s time left over. We’d rather find out first than read about it later
- Talk to customers to validate new solutions and get feedback
- ➡️ Our tech stack: TypeScript, Python, langfuse, PostHog, Google Cloud Platform, mastra, Postgres, MongoDB, and several AI providers.
Benefits
- Versátil working hours- ? You’ve worked with an eval system to measure and improve model performance, rather than a handful of test prompts and a good feeling
- ? You combine engineering rigor with product thinking. You design, debug, and improve end-to-end systems that solve real problems
- ? You’re truly AI-native: you follow new models, papers, and techniques, and you continuously refine how you use them in practice
- ? You’re also AI-native in how you build: you use modern AI-powered dev tools (Claude Code, Cursor, Codex or similar) and assemble your own agents and workflows to move faster
- ? 2+ years of hands-on experience with GenAI and 3+ years of engineering experience overall. One year of very intense AI work also counts, if it’s backed by solid software or data science experience
- ? Fluent in Spanish and comfortable in English
- ? You’ve run GenAI in production at real scale.
As a reference: an agent or workflow doing at least 5,000 executions a month, and you were the one accountable for it. Experiments don’t count
- AI observability platforms like langfuse (tracing, metrics, evals for LLM apps and agents)
- Designing and running evaluations
- Multimodal LLMs, especially vision models for OCR document understanding
- Building and orchestrating multi-agent systems
- RAG systems and vector databases
- Designing or working with MCPs- We move fast and communicate openly. Our goal is to wrap up the process within two weeks
- Intro call with Enric (AI Lead), 25 mins
- Call with Pol (co-founder, CPO), 25 mins
- Technical interviews with Enric and Guillermo (Head of Tech), 2 hours
- Final call with Arnau (co-founder, CEO), 25 mins
📌 AI Engineer (Product) (Barcelona)
🏢 haddock
📍 Barcelona