28 sep
|
Johnson u0026 Johnson
|
Madrid
28 sep
Johnson u0026 Johnson
Madrid
Experteer Overview
As Senior Scientist in the Generative AI Evaluation u0026amp; Standards (EQS) team, you design, build, and run evaluation methods to ensure AI systems for pharma Ru0026amp;D are quality, traceable, and scientifically defensible. You author rubrics, curate benchmark datasets, validate AI judges, and produce readouts that inform release decisions. You work with clinical and regulatory partners to refine evaluation criteria and drive responsible AI adoption. This role offers impact at a leading healthcare innovator, shaping how AI supports scientific work.
Compensaciones / Incentivos
• Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy
• Author evaluation rubrics and scoring criteria; curate golden and synthetic datasets with domain experts; maintain registry of reusable evaluation assets
• Validate AI judges against human expert agreement and run model, prompt, retriever, and agent benchmarks that produce standardized quality readouts
• Analyze failure patterns (hallucination, unsupported claims, weak traceability) and translate findings into actionable recommendations
• Develop therapeutic-area-specific evaluation criteria with scientific, clinical, and regulatory partners; refine them using real-world feedback
• Design evaluation methods for scientific reasoning, evidence synthesis, and hypothesis quality where generic benchmarks fall short
• Build evaluation tooling and reusable patterns to enable self-serve across teams
Responsabilidades
• Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Bioinformatics, Biomedical Engineering, Applied Mathematics, Biostatistics, or related field; PhD preferred
• 6+ years hands-on experience in AI/ML evaluation or data science (Master's) or 3+ years with PhD
• Experience designing and running evaluation frameworks, scientific benchmarks, or quality assessments for AI/ML systems
• Hands-on work with generative AI, LLMs, retrieval-augmented generation, agentic frameworks, and prompt engineering
• Proficiency in Python and modern AI/ML tooling (evaluation harnesses, embedding models, vector databases, LLM APIs)
• Ability to translate expert scientific judgment into measurable criteria and reproducible protocols
• Collaborative, self-driven, able to work across multidisciplinary teams
Requisitos principales
• annual bonus
• vacation days
• parental leave
• well-being reimbursement
• insurance plans
• service anniversary and recognition awards
📌 Senior Scientist - GenAI Evaluation (Madrid)
🏢 Johnson u0026 Johnson
📍 Madrid