16 ago
|
dsm-firmenich
|
Cataluña
16 ago
dsm-firmenich
Cataluña
Job Title: Lead Data Engineer (Cross Domain)Location: Barcelona, SpainJoin us as a Lead Data Engineer to design and maintain robust data pipelines, drive the best deployment practices, and manage key initiatives in the Procurement domain. You will own the full lifecycle of data products — from source ingestion to Gold layer consumption — using modern DevOps practices across Azure DevOps and GitHub.You will collaborate with data modelers, data stewards, and business partners to deliver impactful, production-grade data solutions. Elevate your career by mastering cutting-edge tools — dbt, Databricks, PySpark, and CI/CD automation — in a dynamic, innovative, and general environment.Your key responsibilitiesDesign and implement end-to-end data pipelines for ingestion, transformation, and storage across Bronze (raw), Silver (Raw Vault and Business Vault), and Gold (data marts and semantic layer), ensuring scalable and reliable data processing.Develop and manage data ingestion processes from source systems through ETL and API-based extraction methods where platform tables are unavailable, ensuring seamless integration of enterprise data sources.Build modular, reusable, and maintainable transformation frameworks using dbt, PySpark, and SQL, including the creation of staging models (raw_stage and stage) with hash keys, hash diffs, business keys, and load metadata in alignment with Data Vault 2.0 and modern engineering best practices.Design, build, and maintain CI/CD and DevOps processes in Azure DevOps and GitHub Actions, incorporating automated testing, linting, deployment gates, environment promotion, and standardized branching, pull request, code review, and merge strategies.Implement and manage Infrastructure-as-Code (IaC) and deployment automation for data platform resources, including provisioning environments, deploying dbt jobs and Databricks workflows, configuring clusters, managing dependencies, scheduling jobs,
and maintaining environment-specific configurations.Establish proactive monitoring, observability, and operational excellence practices by implementing alerting, conducting regular pipeline maintenance and upgrades, and troubleshooting issues to ensure platform reliability, performance, and rapid incident resolution.Drive data quality, governance, and compliance standards through automated dbt testing, validation frameworks, FAIR data principles, lineage management, auditability, and data discoverability across all layers of the platform.Provide technical leadership and cross-functional collaboration by evaluating solution feasibility, leading workload planning, mentoring data engineering teams, defining engineering standards, and partnering with data engineers, modelers, stewards, BI developers, data scientists, and business SMEs to deliver impactful data solutions.We offerUnique career paths across health, nutrition, and beauty — explore what drives you and get the support to make it happen.A science-led company with cutting-edge research and creativity everywhere — from biotech breakthroughs to sustainability game-changers, you'll work on what's next.You bringStrong technical expertise in dbt, SQL, Python, Spark (PySpark), Databricks, Git, Azure DevOps, and GitHub, with proven experience designing, building, and maintaining production-grade data pipelines for 5+ years.Deep CI/CD and automation experience, including the design and implementation of deployment pipelines using Azure DevOps Pipelines and GitHub Actions, with Infrastructure-as-Code knowledge (Terraform, Bicep)
considered a strong advantage.Advanced Git and repository management skills, including branching strategies, pull request standards, code review best practices, and governance across Azure DevOps and GitHub environments.Strong cloud data engineering expertise (Azure preferred), with hands-on experience deploying, managing, and operating Databricks jobs, clusters, and workflows in production environments.Extensive knowledge of Data Vault 2.0 architecture, including Raw Vault (Hubs, Links, Satellites), Business Vault components (bSAT, PIT, Bridge), and Gold-layer data marts and semantic models.Expertise in data ingestion and integration patterns, including batch ETL, incremental processing, Change Data Capture (CDC), and API-based ingestion from a wide range of source systems.Strong focus on data governance and quality, with experience applying FAIR data principles, implementing data quality automation, and managing enterprise data platforms; exposure to scientific datasets such as cheminformatics, bioinformatics, or microbiome data is an added advantage.Leadership-oriented and collaborative mindset, combining curiosity, continuous improvement, accountability, and end-to-end ownership with the ability to lead teams, oversee projects, engage stakeholders, and work effectively across functions, geographies, and diverse teams.Join our global team powered by science, creativity, and a shared purpose: to bring progress to life.Whether it’s fragrance that helps you focus, alternative meat that’s better for the planet, or reducing sugar without losing flavor, this is where you help shape the future of nutrition, health, and beauty for everyone, everywhere.And if you have a disability or need any support through the application process, we’re here to help - just let us know what you need, and we’ll do everything we can to make it work.Department:IT
📌 Lead Data Engineer (Cross Domain) (Cataluña)
🏢 dsm-firmenich
📍 Cataluña