Responsibilities
- Evaluate AI-generated coding interactions end-to-end.
- Judge whether agent outputs are useful, broadly correct, and aligned with strong engineering judgment.
- Assess the quality of explanations, reasoning, preambles, and developer guidance.
- Distinguish different levels of response quality and provide clear, opinionated feedback.
- Identify what worked, what failed, and what felt misleading or off.
- Help define what high-quality interactions look like for tools such as Cursor.
Requirements
- Staff- or Principal-level software engineering experience, or equivalent experience.
- Strong background in TypeScript/JavaScript or Python.
- Hands-on experience using OpenAI Codex, Claude Code, and modern AI-assisted development workflows.
- Ability to evaluate code without fully executing it or reviewing every line deeply.
- Comfort making subjective but rigorous judgments and giving direct, opinionated feedback.
- High standards for engineering quality and judgment.
Benefits
- Contract engagement at approximately 10–20 hours per week.
- Scheduled from ASAP through early May, with a possible extension.
- Selection process includes a take-home evaluation exercise and one behavioral interview.
- Nice-to-have experience includes AI-first IDEs such as Cursor, prompt design or evaluation workflows, mentoring senior engineers, or defining engineering standards.
📌 Go Engineer, AI Code Reviewer (España)
🏢 G2i
📍 España