14 sep
|
Jobtailor
|
Madrid
- Validate high-value emerging AI and automation technologies and de-risk their adoption across dLocal
- Own technology scouting, prototyping, and evaluation for dLocal
- Run instrumented spikes and benchmarks on LLMs, agentic systems, vector databases, orchestration frameworks, copilots, assistants, and other emerging technologies
- Compare vendor and open-source options across quality, cost, latency, security, and integration complexity
- Deliver decision memos with recommendations to adopt, watch, or avoid
- Design and maintain evaluation environments with datasets, prompts, scenarios, and telemetry
- Build automation and tooling to measure quality, robustness, latency, cost, and regressions
- Build focused prototypes to explore architecture, integration patterns, operational constraints, security boundaries, and failure modes
- Define readiness guidance, patterns, guardrails, limitations, operational considerations, and integration requirements
- Coordinate hand-offs to engineering teams responsible for productionization and support transitions as needed
- Track outcomes of Lab recommendations to improve evaluation methods
- Work with Security, Legal, Compliance, and AI teams on risk assessments and governance recommendations
- Maintain reusable checklists, decision templates, and standards
- Incorporate learnings from external copilots and the AWS AI suite into adoption guidelines
- Partner with AI and domain teams to ensure collaboration and clear boundaries
- Participate in hiring as a technical evaluator and culture champion
- Mentor engineers on evaluation methods, benchmarking, and experimental design
- Share knowledge through internal write-ups, tech talks, meetups, and conferences
Requirements
- 8+ years of software engineering experience, including significant experience operating at senior or Staff-level scope
- Deep hands-on experience building and evaluating systems based on LLMs and modern AI tooling
- Strong software engineering fundamentals and ability to rapidly build high-quality experimental systems
- Experience building agentic or multi-step AI systems involving tool use, orchestration, state, retrieval, or external integrations
- Strong knowledge of cloud infrastructure, preferably AWS, and ability to run experimental workloads securely and cost-consciously
- Experience with observability, telemetry, testing, and benchmarking of complex systems
- Ability to reason about system architecture, reliability, scalability, asynchronous workflows, and distributed components
- Track record of designing experiments or benchmarks that influenced meaningful technical decisions
- Experience constructing evaluation datasets, including task selection, labelling, and holdout discipline
- Working knowledge of LLM-as-judge methods, human evaluation, inter-annotator agreement, and their appropriate use
- Ability to reason about statistical significance on small samples
- Familiarity with regression tracking, telemetry, and versioning
- Ability to turn ambiguous ideas into scoped evaluation plans with hypotheses and metrics
- Comfortable making trade-off calls across quality, latency, cost, and vendor lock-in
- Experience writing concise decision memos
- Ability to explain technical results to non-specialists
- Experience working with platform, product, and operations teams
- Ability to influence without authority and align teams around standards and guardrails
- Curious, experimentation-oriented mindset with disciplined measurement and risk awareness
- Comfortable in a small, high-leverage team without embedded PMs
- Builder attitude favoring reusable tools, templates, and playbooks
Core Competencies
Demonstrates extensive experience in evaluating and adopting AI and automation technologies, with a strong focus on LLMs and cloud infrastructure, particularly AWS. Capable of designing experiments, building prototypes, and collaborating across teams to ensure effective integration and governance.
Highest-signal resume keywords
- LLM Evaluation
- Cloud Infrastructure (AWS)
- Experimental Design
- Benchmarking and Telemetry
- Decision Memo Writing
Hard Skills
- Software Engineering
- AI Tooling
- System Architecture
- Observability
- Statistical Significance Reasoning
- Evaluation Dataset Construction
- Regression Tracking
- Integration Patterns
- Automation and Tooling
- Prototyping
Soft Skills
- Influencing Without Authority
- Curiosity
- Collaboration
- Mentoring
- Communication
Industry Keywords
- AI Governance
- Risk Assessment
- Compliance
- Experimental Workloads
- Vendor Evaluation
Tools & Technologies
- LLMs
- Agentic Systems
- Vector Databases
- Orchestration Frameworks
- Telemetry Tools
- Decision Templates
- Reusable Checklists
- AWS AI Suite
#J-18808-Ljbffr
📌 Staff AI Engineer – AI Labs (Madrid)
🏢 Jobtailor
📍 Madrid