06 sep
|
Gangkhar-ES
|
Madrid
06 sep
Gangkhar-ES
Madrid
At Gangkhar, we’re building the next-generation insurance infrastructure. Our AI-native protection platform enables partners to design, deploy, and scale world-class protection programs in just a few weeks.
We're looking for an AI Engineer with a hands-on mindset and a product mentality. You'll build the agent platform that powers Gangkhar: the infrastructure where AI agents are designed, evaluated, governed, and operated at scale. You'll work on Sherpa Mesh, our internal reference agent platform built and maintained by our infrastructure team — extending it, building on it, and, when needed, contributing to it directly. You'll collaborate closely with the architects who own client discovery and agent design, turning their specs into production-grade agents.
What Kind of Engineer We’re Looking For
This is a role for an engineer who cares how the code is built, not only whether it runs.
- You build capabilities, not one-offs. Faced with a stakeholder-specific request, you find the reusable shape underneath it — and you know when a request genuinely is specific.
- You think in modules and boundaries. You know what belongs together, what doesn't, and you can say why.
- You design before you type, and you can defend a design in a conversation with an architect and in plain language with a non-technical stakeholder.
- You are precise: clear names, explicit behaviour, no guessing at what a function does from the outside.
- You leave a codebase more coherent than you found it, and you read an unfamiliar system with its grain before proposing changes.
- You work with coding agents daily and own every line they produce. Output volume is free now; judgment is the scarce part — we want engineers who reject their agent's work, not who ship it.
- If "it works for this client, ship it" is your standard, this isn’t the role.
Your Impact
- Design,
build, and deploy LLM-powered agents and multi-agent systems within Sherpa Mesh, our internal agent platform (agent manifests, registry, runtime, delegation, fleet coordination).
- Build directly on LLM APIs served through Azure AI Foundry and OpenRouter: agent loop, tool calling, context engineering, without heavyweight orchestration frameworks.
- Extend and operate the agent memory pipeline — extraction, property injection, retrieval — within the existing attribute/property/memory architecture.
- Take the evaluation harness from early-stage production signal detection to a real offline eval suite: datasets, graders, regression tests, and the promotion gate that decides what goes to production.
- Implement observability for agentic systems: run-level tracing, token accounting, debugging tools.
- Apply guardrails and governance: attribute-based access policies, PII handling, human-in-the-loop flows.
- Integrate agents with internal APIs and business systems via open protocols (MCP) to trigger real-world actions.
- Make pragmatic engineering trade-offs between speed, quality, and scalability.
What You Bring
- 5+ years building and operating backend systems. Deep, not broad-and-shallow — plus 1–2 years building LLM-based agents or GenAI systems in production.
- Strong TypeScript/Node.js, and the judgment to use the type system rather than fight it. Real production ownership of PostgreSQL, HTTP API design,
job queues— not just familiarity.
- Comfortable reading and writing Python — not your main language, but you'll touch it.
- Experience building agents directly against LLM APIs, and the judgment to explain why you didn't reach for a framework.
- Judgment about context engineering, tool design, and retrieval — the interesting problems are in the interfaces, not the prompts.
- Experience with retrieval architectures: RAG pipelines, knowledge base construction, and general understanding of graph-based retrieval (GraphRAG, knowledge graphs).
- Experience with evals and LLM observability (eval harnesses, tracing, quality metrics).
- Deployment with Docker and Kubernetes; cloud experience (Azure preferred).
- Awareness of security and compliance: GDPR, PII masking, access control, AI safety mechanisms.
- Product mentality: you understand the business logic behind what you're building, not just the spec. When an architect's design has a gap or doesn't hold up in practice, you push back with a better alternative — you don't build it blind and let it fail downstream.
- Familiarity with open agent interoperability protocols (A2A, Agent Cards) is a plus.
- Cost and latency reasoning — you can estimate token budgets and per-query costs, and know when to route to a cheaper or faster model instead of defaulting to the biggest one
Tech Stack
TypeScript on Node.js, with Hono. PostgreSQL for storage, background jobs via a job queue. Model access through Azure AI Foundry and OpenRouter. MCP for agent interoperability. Deployed on Docker/Kubernetes in Azure. Python for evaluation tooling. Go and Preact exist in the codebase (CLI, internal devtools) but sit with the infrastructure team, not day-to-day for this role.
Languages
Spanish and Fluent in English (required)
📌 Artificial Intelligence Engineer (Madrid)
🏢 Gangkhar-ES
📍 Madrid