pMindrift is building a dataset to evaluate AI coding agents by creating realistic developer environments, tasks, and tests. You design prompts, define success criteria, and ensure solutions are verifiable across multiple valid approaches. /ppYou will iteratively refine tasks based on QA feedback, aiming for fair evaluation of models while managing complexity and realism in simulated environments. /p #J-18808-Ljbffr