Responsibilities
- Evaluate AI-generated coding interactions from end to end.
- Judge whether outputs are useful, correct at a high level, and aligned with strong engineering practice.
- Assess the quality of explanations, preambles, and reasoning, not only the code.
- Distinguish between different levels of response quality.
- Provide clear, opinionated written feedback about what worked, failed, or felt misleading.
- Help define what high-quality interactions look like for Cursor, Codex, and Claude Code.
- Record short video explanations of evaluation feedback.
Requirements
- Senior, staff, or principal-level software engineering experience, or equivalent experience.
- Strong background in TypeScript/JavaScript or Python.
- Hands-on experience with at least one of OpenAI Codex, Claude Code, or Cursor.
- Deep familiarity with modern AI-assisted development workflows.
- Ability to evaluate code without executing it or reviewing every line.
- Strong written and spoken English at B2 level or above.
- Comfort giving direct, opinionated feedback and maintaining a high engineering quality bar.
Benefits
- Remote worldwide contract engagement.
- Adaptable schedule of 10–20 hours per week.
- Ongoing projects generally lasting about two weeks to a few months, with new projects offered to successful evaluators.
- Start as soon as the take-home exercise is completed and a project has an open seat.
- Hiring process consists of one take-home evaluation exercise with a recorded Loom walkthrough and no interview.
📌 Senior Software Engineer (España)
🏢 G2i
📍 España