AI Benchmark Engineer | Native Language Specialist - Spanish (Spain) - Remote (Madrid)

AI Benchmark Engineer | Native Language Specialist - Spanish (Spain) - Remote (Madrid)

04 oct
|
LILT (Production)
|
Madrid

04 oct

LILT (Production)

Madrid

We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks.

You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches. Note this is a remote, freelance opportunity What You'll Deliver Task Engineering: Evaluating Coding Agents. Crucially, these assets must remain in the target language to genuinely measure multilingual handling.

Prompting & Translation: finding failure points where AI does not work, in your native language Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).

Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.

Experience: 5+ years of industry experience in software engineering.

Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities. High English proficiency.

Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing. Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.

Domain Expertise:



Deep technical understanding of multilingual text processing pitfalls, including: For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts. As an independent contractor, work when you want, as much or as little as you want. Work on projects that actually matter.

Contribute to cutting-edge AI and language technology that is shaping how humans and machines communicate. Join a general community of linguists, subject matter experts, and language professionals who are advancing human knowledge together. As a Lilt contractor you get access to diverse, innovative projects that expand your portfolio and sharpen your skills across industries and domains.

Bring your language skills to life on projects that are as interesting as they are impactful. We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

Not ideal as a full time job or primary income source. Work availability fluctuates with project demand, making this better suited as a supplemental income stream. As a 1099 contractor, you won't receive benefits such as health insurance, paid time off,



or retirement contributions, and hours are not guaranteed.

Once you accept a task, we expect quality work and on-time delivery. We cannot engage contractors in regions subject to international embargo or sanctions. As a 1099 contractor, you are solely responsible for your own tax obligations.

We recommend consulting a tax professional before engaging. 1 - Submit your application including an updated copy of your CV in English 3 - Finalize onboarding and profile set-up in our system, and become eligible for Applied AI projects. Join our global community who thrive on innovation and excellence. Our collective knowledge, uniqueness, and skills deliver multilingual AI and human-verified services to Enterprises, Governments, and AI Developers around the world.

Work on diverse projects from anywhere, any time you want. Get paid quickly and fairly, and build your professional network in a supportive community—all through a streamlined application process tailored to your expertise. As part of our recruitment efforts, we may use artificial intelligence (AI) and automated tools to assist in the evaluation of applications, including résumé screening, assessment scoring, and interview analysis.

These tools are designed to support human decision-making and help us identify qualified candidates efficiently and objectively. We extend equal opportunity to all individuals without regard to an individual's race, religion, color, national origin, ancestry, sex, sexual orientation, gender identity, age, physical or mental disability, medical condition, genetic characteristics, veteran or marital status, pregnancy, or any other classification protected by applicable local, state or federal laws.

📌 AI Benchmark Engineer | Native Language Specialist - Spanish (Spain) - Remote (Madrid)
🏢 LILT (Production)
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai benchmark engineer | native language specialist - spanish (spain) - remote (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai benchmark engineer | native language specialist - spanish (spain) - remote (madrid) / madrid