AI Benchmark Engineer | Native Language Specialist (Madrid)

AI Benchmark Engineer | Native Language Specialist (Madrid)

04 oct
|
LILT (Production)
|
Madrid

04 oct

LILT (Production)

Madrid

We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.
We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches.
Note this is a remote, freelance opportunity
What You'll Deliver Task Engineering: Evaluating Coding Agents.
Crucially, these assets must remain in the target language to genuinely measure multilingual handling.
Prompting & Translation: finding failure points where AI does not work, in your native language
Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).




Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.
Experience: 5+ years of industry experience in software engineering.
Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.
High English proficiency.
Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing.
Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.
Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including:
For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts.
As an independen

📌 AI Benchmark Engineer | Native Language Specialist (Madrid)
🏢 LILT (Production)
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai benchmark engineer | native language specialist (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai benchmark engineer | native language specialist (madrid) / madrid