11 ago
|
Intellias
|
España
Senior AI/RAG Engineer (Document Intelligence)
Location: Remote from Spain (an indefinite Spanish employment contract) Our client is a leading general investment management company headquartered in London. It manages over $228 billion in assets and serves institutional investors, pension funds, wealth managers, and other sophisticated clients worldwide. The firm specializes in quantitative investing, alternative investments, systematic trading strategies, and technology-driven asset management.
Data science, machine learning, and AI are core components of its investment and research processes.
As part of our collaboration we will focus on two foundational capabilities required to enable safe and scalable AI adoption across the enterprise: Agentic Security and AI-Ready Data Foundations.
Project Overview
We build the data foundations that make AI useful and safe inside regulated financial firms. The value of AI is capped by the data its agents can reach: if an agent cannot find, interpret, trace or be correctly permissioned against data, the capability is useless, or worse, unsafe. Your job is to close that gap.
This is a hands-on senior role for an excellent Python engineer with strong data-engineering skills who is genuinely comfortable building with AI agents. You will design and build the catalogue, semantic, entitlement and analytical layers that turn large on-premise data estates into something agents can use.
Requirements
5+ years of production Python development, including 2+ years of LLM and RAG engineering in production: retrieval pipelines, vector stores, structured extraction, and the surrounding operational tooling.
Strong experience building evaluation harnesses: versioned test suites derived from real question banks, separate scoring of retrieval and answers, scoring where a correct "not found" counts as a pass, and suites wired into delivery as release gates (e.g. Langfuse, RAGAS, DeepEval or similar, plus custom metrics).
Hybrid retrieval engineering: keyword and semantic search combined, result merging, cross-encoder reranking, tuning against measured baselines, working within a platform-fixed embedding model and index.
Structure-aware document processing: layout-aware parsing and chunking that keeps tables intact (e.g. Docling,
Tika or similar), including OCR handling for scanned documents and multilingual content.
LLM extraction at scale: schema-driven extraction of attributes, entities, clauses and relationships with per-field confidence, calibrated thresholds and a human review loop (e.g.
Label
Studio or similar), piloted and measured before scale-out.
Strong PostgreSQL: typed relational modelling plus JSONB, schema-as-code with migration tooling, derived views managed as tested transformations.
Provenance and citation discipline: every extracted fact traceable to its source document and passage; answers that state explicitly when something could not be confirmed.
Fluent English for written and spoken communication with client teams.
Will be a plus
Integration with SharePoint and Microsoft Graph APIs or an equivalent enterprise content platform: change notifications, delta polling, permission metadata.
Bitemporal modelling (execution versus effective dates, as-of queries) and document lineage or supersession modelling.
Graph engines (e.g. Apache AGE, Neo4j or similar); controlled knowledge graphs with provenance on every element.
Legal, contract or fund-documentation domain knowledge, or demonstrated ability to learn a document domain in depth.
Permission-aware retrieval: access lists stored in the index, group resolution at query time, denial on uncertainty (the entitlement design is owned by a parallel workstream; this role implements against it).
Token-efficiency engineering: corpus preparation, section slicing, cost measurement per query.
Exposing capabilities to an assistant platform as MCP tools or skills; day-to-day use of AI coding agents.
Experience in financial services or other regulated on-premise environments; client-facing experience.
Responsibilities
Build the evaluation foundation at the start:
convert the client's existing question bank and diagnostic results into a versioned automated test suite with expert-confirmed outcomes, establish the baseline, and mine historical support records as a second ground-truth set.
Improve the client's existing RAG service in place: hybrid retrieval with reranking, structure-aware chunks, context notes, early metadata filters, reliability monitoring, with every change measured against the suite before release.
Build the metadata and entity extraction pipeline under a governed three-tier schema (universal envelope, versioned per-collection specifications, open discovery tier), with a stratified pilot, confidence-routed human review and coverage dashboards.
Build the synchronisation plane over the client's content platform (change notifications, delta polling, scheduled full reviews) shared by retrieval, extraction and downstream views.
Build document lineage and temporal views: supersession and amendment chains extracted only where stated in the text, a bitemporal effective-terms view, derived document status, and a queryable obligations register, with human verification for high-stakes chains.
Contribute to the controlled knowledge graph and query orchestration: typed nodes and edges carrying provenance and confidence, routing between structured lookup, filtered retrieval and lineage views, explicit completeness statements, and typed gaps reported as answers.
Operate the delivered capabilities: releases gated on evaluation results, freshness bounds enforced by withholding stale data, coverage and quality dashboards, corpus health checks reported to document owners.
Why this position
This role sits at the intersection of data engineering, AI, and financial services, solving one of the most important challenges in enterprise AI: enabling agents to securely access and reason over trusted data. You'll have the opportunity to design and build foundational platforms that combine large-scale data systems, governance, and AI technologies in highly regulated environments. It offers significant technical ownership, exposure to cutting-edge AI agent architectures, and the chance to shape how organisations safely unlock value from their data.
📌 Senior AI/RAG Engineer (Document Intelligence) (España)
🏢 Intellias
📍 España