12 sep
|
Jobtailor
|
Madrid
Keep production Databricks pipelines running, including ingestion, transformation, and delivery to downstream consumers
Diagnose and resolve pipeline failures and data quality issues
Reverse-engineer and document existing transformation logic and business rules
Migrate legacy tables from Hive Metastore to Unity Catalog
Maintain Iceberg-enabled table sharing between Databricks and Snowflake
Build and test dbt models, including incremental materializations and data tests
Develop and maintain Airflow DAGs for orchestration
Validate migrated pipelines against Databricks outputs
Contribute to Snowflake modeling, performance, and cost decisions
Work directly with client stakeholders on technical topics alongside the team lead
Requirements
3+ years operating production data pipelines
PySpark and SQL — able to read, debug, and modify existing pipelines
Delta Lake: MERGE/upsert patterns, table properties, OPTIMIZE, partitioning
Databricks Workflows, cluster configuration, job troubleshooting
Unity Catalog: catalogs, schemas, grants, lineage, and the metastore model
Snowflake warehouses, roles and grants, and general operating model
Snowflake query performance and awareness of compute cost behavior
dbt models, sources, tests, and incremental materializations
dbt project structure and deployment workflow
Airflow DAGs, operators, scheduling, and dependency management
Airflow retries, backfills, and idempotent task design
Strong SQL, including window functions,
complex joins, and reading transformation logic
Python for scripting, automation, and API integration
Incremental loading patterns, idempotency, late-arriving data, and reprocessing
AWS S3 and IAM basics
Basic working knowledge of Redshift and its role in wider architecture
Fluent English
Self-directed and able to progress on an unfamiliar codebase without structured onboarding
Able to explain production incidents to non-technical stakeholders and provide realistic ETAs
Core Competencies Demonstrates expertise in managing production data pipelines using Databricks, including ingestion, transformation, and delivery processes. Proficient in SQL, PySpark, and dbt for building and testing data models, with a strong understanding of Snowflake and Airflow for orchestration and performance optimization.
Highest-signal resume keywords
Databricks Pipeline Management
SQL Proficiency
PySpark Development
Airflow DAG Development
Dbt Model Building
Hard Skills
SQL
PySpark
Dbt
Airflow
Delta Lake
Unity Catalog
Snowflake
AWS S3
Incremental Loading Patterns
Data Quality Diagnosis
Soft Skills
Self-Directed
Effective Communication
Stakeholder Engagement
Industry Keywords
Data Pipeline
Data Transformation
Data Quality
Data Modeling
Orchestration
Tools & Technologies
Databricks
Snowflake
Airflow
Hive Metastore
Redshift
#J-18808-Ljbffr
📌 Mid Data Engineer (Madrid)
🏢 Jobtailor
📍 Madrid