About the Role * Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. * Contributors help evaluate and improve frontier AI coding models through structured technical assessments. * The work focuses on realistic infrastructure engineering workflows and model evaluation. * Spots are limited and filling quickly on a first come, first serve basis. What You'll Do * Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. * Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation. * Identify bugs, edge cases, reliability issues, and failure modes. * Compare outputs from multiple frontier models and assess their strengths and weaknesses. * Apply professional engineering judgment to realistic infrastructure engineering scenarios.
Time Commitment * Sprint based project that runs in 12-24 hour stretches based on client requirement. Compensation * $400 per accepted task. * Typical tasks take approximately 2–3 hours after ramp-up. * Compensation is tied to accepted work. Who Should Apply * 2+ years of professional DevOps, SRE, or Cloud Engineering experience. * Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling. * Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. * Ability to evaluate model-generated infrastructure and reliability engineering solutions. * Experience supporting production-scale systems is preferred. #J-18808-Ljbffr
📌 DevOps Engineer - AI Model Evaluator (Madrid)
🏢 Mercor
📍 Madrid