17 sep
|
Farmatalento
|
Madrid
17 sep
Farmatalento
Madrid
We are looking for a Lead Data & Platform Engineer to join the Health Data Platform and Innovation Unit of a Spanish healthcare Consulting, Research and Innovation company, founded in 2022 and currently undergoing significant growth.
Technical ownership of the Data Platform: Architecture, Data Model, Source Integration, Continuous Ingestion, Quality, Performance and Operations. This is not a maintenance role, nor one focused on implementing decisions made by others: the person in this position will define the architecture, document it and take ownership of its long-term operation and evolution. Consolidate and scale the Clinical Data Platform.
Redesign the Data Model, Performance and Traceability to support growth in data volume, data sources and the number of simultaneous projects. Integrate new Healthcare Data Sources. Incorporate new Clinical, Hospital and Administrative Data Sources into the common model, addressing the heterogeneity of formats, terminologies and identifiers inherent to the healthcare sector.
Design and operate Continuous Ingestion and Processing of Healthcare Data using Event-Driven Architectures, with HL7 FHIR Interoperability and under strict data security and sovereignty requirements. Define the Unit's Engineering practices, document the Architecture and technical Decisions, and participate in building and developing the team as it grows. You will work closely with Data Analysts, Data Scientists, Biostatisticians and the Scientific and Clinical Teams to ensure that data is available, validated, traceable and ready for use across projects.
On-premise PostgreSQL (proprietary clinical data model): tuning, partitioning, RLS and replication. Python for data processing, integration and automation.
Apache Kafka / Confluent Platform: event ingestion and processing.
Change Data
Capture from healthcare systems (Debezium, Kafka Connect).
Kafka
Streams, Apache Flink or ksqlDB.
Schema governance and data contracts: Schema Registry, Avro/Protobuf, versioning and compatibility.
Common Data Models: OMOP CDM and other applicable European standards. Orchestration and version-controlled transformation: Airflow or Dagster, dbt. Kubernetes, Docker, CI/CD, and on-premise and hybrid-cloud deployments. LLM and agent pipelines operating on structured clinical data.
Implement Change Data
Capture and in-transit validation for healthcare data.
Design the interoperability layer: transformation into FHIR R4 resources, profiles and Implementation Guides, and their representation as data contracts. Translate clinical and regulatory requirements into verifiable architectural decisions covering latency, traceability, auditability and data quality in transit. Work with clinical teams to translate clinical validation and plausibility rules into executable logic across data flows.
Platform and Data Model Redesign and optimise the PostgreSQL relational model: partitioning, indexing, study/user-level RLS, tuning and replication. Design, manage and optimise databases for clinical and Real-World Data, ensuring quality, integrity, traceability and security throughout the entire data flow. Define the data architecture required to support the company's growth in terms of sources, volume and clients.
Integration of New Data Sources Build extraction pipelines from EHR systems, including structured and unstructured data, clinical records and administrative data.
Incorporate new sources into the common model: source analysis, mapping, patient identifier resolution and data-load quality control. Implement and maintain ETL/ELT processes for integrating, transforming and loading data from multiple sources. Curate and normalise data into CDMs (Common Data Models).
Map data to clinical terminologies including SNOMED CT, ICD-10, ATC and LOINC. Security, Compliance and Data Sovereignty Design and operate the platform in compliance with ENS High Category requirements and GDPR/LOPDGDD requirements applicable to health data. project-level segregation; Support on-premise or network-isolated deployments when required by data sovereignty frameworks. Reproducibility and Quality Implement automated data-quality testing and pipeline observability.
Implement and maintain Common Data Models (CDMs), such as OMOP and other applicable European standards. Develop and maintain MCP servers for controlled access to clinical data, including authentication, project-level scoping and query auditing.
Technical
Leadership, DevOps and Team Establish the Unit's engineering standards: code reviews, documented Architecture Decision Records (ADRs) and technical debt management.
Apply DevOps practices: Docker, Kubernetes, CI/CD, Infrastructure as Code, and management of on-premise and hybrid-cloud deployments. Lead or participate in technical discussions with internal and external IT teams and technology vendors. Participate in the recruitment, onboarding and development of the Data Engineering team.
Prepare and maintain technical documentation covering processes, databases and workflows. Bachelor's Degree in Computer Engineering/Computer Science, Telecommunications Engineering, Mathematics or a related discipline. A Master's Degree or postgraduate studies in Big Data, Databases or Artificial Intelligence will be highly valued.
Minimum of 10 years' experience in Data Engineering or Data Platform Architecture,
including at least 3–5 years as a Technical Lead, Lead Engineer or Architect. Proven experience leading the end-to-end development or modernisation of a mission-critical data platform at regional, national or corporate scale, with simultaneous responsibility for architecture, implementation and production operations.
Experience in regulated environments or the public sector (Healthcare, Public Administration, Banking, Defence, Energy), involving demanding requirements for security, traceability and data sovereignty.
Experience acting as the technical point of contact with technology vendors and public-sector organisations, including architecture committees, project monitoring, technical specifications and public tender documentation.
Experience working with real-world clinical data (EHRs, registries and hospital data).
Apache Kafka / Confluent Platform in production: topic design and partitioning, Kafka Connect, Schema Registry, delivery semantics, multi-cluster replication, security (mTLS, RBAC, ACLs) and day-to-day operations.
Official
Confluent certification (Kafka Developer and/or Kafka Administrator), or demonstrably equivalent expertise, will be highly valued.
Change Data
Capture using Debezium or equivalent technologies over relational data sources. Distributed event processing using Kafka Streams, Apache Flink or ksqlDB.
Advanced PostgreSQL: tuning, EXPLAIN, partitioning, RLS and replication.
Advanced
Python for data processing and automation, together with sufficient proficiency in the JVM ecosystem (Java/Scala) used by Kafka/Flink. Kubernetes and deployment of distributed systems on-premise and in hybrid-cloud environments; Docker, CI/CD and Infrastructure as Code.
Observability and operation of distributed systems: instrumentation, alerting, production incident diagnosis and performance analysis. Pipeline orchestration (Airflow, Prefect or Dagster) and version-controlled transformation (dbt or equivalent). Version control (Git), software engineering best practices and Agile methodologies.
Professional-level English (C1 or above). Interaction with international technology vendors, technical documentation and product support is conducted in English. You define the architecture, document it and defend your decisions.
The data you will be working with is used to inform decisions on treatments and healthcare policies. Continuous ingestion of healthcare data, interoperability, data sovereignty and clinical data quality. Training budget and official certification in the technologies included in the stack.
Compensation aligned with a senior technical leadership profile, to be determined according to the experience and expertise provided. Permanent employment contract. Hybrid working model – 3 days in the office and 2 days working remotely.
📌 Lead/Senior Data Engineer (Madrid)
🏢 Farmatalento
📍 Madrid