Research Data Scientist (Barcelona)

Research Data Scientist (Barcelona)

04 ago
|
Fundamental
|
Barcelona

04 ago

Fundamental

Barcelona

ph3bAbout Fundamental /b /h3 pFundamental is an AI company pioneering the future of enterprise decision-making. Founded by DeepMind alumni, Fundamental has developed NEXUS – the world's most powerful Large Tabular Model (LTM) – purpose-built for the structured records that actually drive enterprise decisions. Backed by world class investors and trusted by Fortune 100 companies, Fundamental unlocks trillions of dollars of value by giving businesses the Power to Predict. /p pAt Fundamental, you'll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world's largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI. /p h3bKey responsibilities /b /h3 pAs part of the Research team, you will contribute to the development of breakthrough machine learning models by working on one of the most important frontiers in model training and evaluation: high-quality real and synthetic data. /p pThis role is especially focused on synthetic data generation, Structural Causal Models (SCMs), and realistic simulation-based data sources. You will help us design, evaluate, and scale datasets that capture the structure, dependencies, and edge cases needed to train foundation models for enterprise tabular data. /p pThe main responsibilities of this role are: /p ul liIdentifying, characterizing, and evaluating high-value data sources for training and evaluating ML models, including real-world data, synthetic data, SCM-generated data, and physical or systems-based simulator outputs /li liDesigning and analysing synthetic data generation approaches based on Structural Causal Models, probabilistic models, simulators, and other mechanisms that capture realistic relationships between variables /li liWorking with researchers to define what makes a synthetic dataset useful,



realistic, diverse, causally meaningful, and appropriate for model training or evaluation /li liBuilding tools and workflows to generate, validate, benchmark, and iterate on synthetic datasets at scale /li liDeveloping metrics and evaluation procedures for synthetic data quality /li liTransforming structured, unstructured, simulated, and causally generated data into formats suitable for training and evaluating large-scale ML models /li liCollaborating with the research team to maintain a reliable, efficient training pipeline where data quality, data diversity, and synthetic data generation are critical components /li liCollaborating with the wider engineering and infrastructure team to ensure data generation and processing workflows are scalable, reproducible, and robust /li /ul h3bMust have /b /h3 pbExperience with: /b /p ul liSynthetic data generation for machine learning, especially for structured or tabular data /li liStructural Causal Models, causal graphs, causal inference, probabilistic modelling, or simulation-based data generation /li liIdentifying and evaluating high-quality data sources to train and evaluate ML models, including both real-world and realistic synthetic data sources /li liBringing data from structured and unstructured sources, simulators, causal models, or generative processes into formats accessible by ML models /li liDesigning quantitative analyses to assess data quality, realism, diversity, bias, coverage, and downstream model performance /li /ul pbStrong fundamentals in:



/b /p ul liStatistics, probability, and applied machine learning /li liData science workflows, including exploratory analysis, feature understanding, validation, and experimental design /li liSoftware engineering for research-grade and production-grade data workflows /li /ul pbStrong knowledge of: /b /p ul liPython data processing and scientific computing stack, including numpy, pandas, scipy, scikit-learn, or similar tools /li /ul pbFamiliarity with: /b /p ul liCausal modelling, graphical models, probabilistic programming, agent-based simulation, discrete-event simulation, or physical / systems-based simulators /li liData storage and data versioning solutions /li liClassical machine learning and deep learning methods, especially outside of purely LLM-based workflows /li /ul h3bNice to have /b /h3 ul liContributions to open source ML, causal inference, synthetic data, simulation, or data science projects /li liBSc, MSc, or PhD in computer science, machine learning, statistics, mathematics, physics, engineering, economics, or another quantitative field /li liExperience working with tabular data, predictive analytics, or enterprise decision-making systems /li liExperience building or evaluating synthetic datasets for model training /li liExperience with SCM libraries, probabilistic programming frameworks, simulation environments, or custom data generation pipelines /li /ul h3bBenefits /b /h3 ul liCompetitive compensation with salary and equity /li liComprehensive health coverage for you and your dependents /li liPaid parental leave for all new parents, inclusive of adoptive and surrogate journeys /li liRelocation support for employees moving to join the team in one of our office locations /li liA mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action /li /ul /p #J-18808-Ljbffr

📌 Research Data Scientist (Barcelona)
🏢 Fundamental
📍 Barcelona

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: research data scientist (barcelona) / barcelona

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: research data scientist (barcelona) / barcelona