30 jul
|
F. Hoffmann-La Roche
|
Madrid
30 jul
F. Hoffmann-La Roche
Madrid
At Roche you can show up as yourself, embraced for the unique qualities you bring.
The Position The Data
Scientist - Enterprise Search role is responsible for contributing to the design and development of the next-generation Enterprise Search and information retrieval architectures. The data scientist will develop advanced AI solutions, with a strong focus on Generative AI, LLM-based applications, and scalable data services. This requires to work with large datasets, develop and evaluate machine learning models, and collaborate with cross-functional teams to improve the accuracy, coverage, relevance, and performance of search algorithms at enterprise scale.
This role involves direct communication with project stakeholders and contributes to team best practices, while identifying optimization opportunities that enhance the impact of moderately complex data solutions within larger product architectures. You will leverage advanced technical skills to translate business needs into actionable data science initiatives.
Job Responsibilities
Generative AI, Agentic AI and LLM Optimization Model Development & Experimentation: Lead exploratory data analysis, feature engineering, model selection, training, validation, and performance evaluation for machine learning and AI-enabled solutions.
Experimentation and Innovation: lead experimental projects and drive innovation in enterprise search, exploring novel approaches like GraphRAG or agentic search patterns. RAG Experimentations (RAG Evaluation Framework): Design and conduct Retrieval-Augmented Generation experiments to evaluate and improve search relevance and performance.
LLM Model Evaluation: Evaluate the performance of Large Language Models in various enterprise search contexts, ensuring they meet business requirements and performance standards. Support data-driven prioritization and strategic decision making through actionable insights, predictive models, and operational intelligence. Conducts A/B testing and experiments to assess the performance of search models and algorithms.
Consultancy
Provide expert consultancy on data science and machine learning best practices, guiding internal teams and stakeholders. Design and lead proof-of-value (PoV) projects, conduct knowledge-sharing sessions, and deliver impactful demos to showcase capabilities and gather feedback. Data Engineering & Vector Databases Data Engineering & Processing: Work with structured and unstructured data, building efficient pipelines for data ingestion, preprocessing, and feature engineering.
Vector database experimentation: Conduct experiments with vector databases to improve the efficiency and accuracy of our search systems.
Data quality: Design data enhancement modules that extract, enrich, and validate document content and metadata, directly improving downstream model context, search recall, and agentic reasoning.
Model
Lifecycle and integration Deployment, Testing, and Training of ML Models and Endpoints: Develop, deploy, and continuously refine machine learning models and endpoints to enhance search functionalities. Conduct rigorous testing and validation to ensure model accuracy and reliability.
MLOps & Monitoring: Implement best practices for model deployment, versioning, monitoring, and performance optimization.
Qualifications Education / Experience Master's degree or PhD in Data Science, Computer Science, Statistics, Mathematics, Engineering, Artificial Intelligence, or a related quantitative field. Demonstrated experience as a rising expert developing predictive models and leading specific analytical modules or project components. Proven track record of taking full accountability for the quality and timely delivery of analytical tasks and troubleshooting complex data issues independently.
Experience working effectively on moderately complex data science problems and understanding how contributions fit into medium-sized data architectures.
Technical Skills
Shows strong proficiency in programming languages, particularly Python. Has proven experience as a Data Scientist, preferably with a focus on information retrieval and NLP. Possesses hands-on experience with machine learning and deep learning frameworks and libraries (e.g., TensorFlow, PyTorch, scikit-learn).
Familiarity with cloud platforms and services, particularly AWS or Microsoft Azure. Familiarity with version control systems (e.g., Git) and agile development practices. Elasticsearch, Solr, or Lucene) and solid understanding of search algorithms, information retrieval, and relevancy tuning is a plus.
Additional Qualifications
Strong communication and collaboration skills, with the ability to manage direct communication with immediate project stakeholders. Ability to actively integrate feedback from technical peers and junior team members. Proactive mindset to identify potential optimizations or new analytical approaches within the project scope.
Together, more than 100'000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a integral impact.
We believe it's urgent to deliver medical solutions right now - even as we develop innovations for the future. We commit ourselves to scientific rigor, unassailable ethics, and access to medical innovations for all. We are proud of who we are, what we do, and how we do it.
We are many, working as one across functions, across companies, and across the world.
📌 Data Scientist (Python / AWS) (Madrid)
🏢 F. Hoffmann-La Roche
📍 Madrid