Junior Data Engineer - España (Madrid)

Junior Data Engineer - España (Madrid)

18 sep
|
Cato
|
Madrid

18 sep

Cato

Madrid

Your mission Bring a tender from the source portal into Cato: scraping, parsing, merging, enrichment.
Youll start by owning a handful of sources end to end - the scraper, the job behind it, and the data that comes out - and take on more as you go.
Not tickets handed to you: sources youre responsible for.
What youll actually do * Build and maintain scrapers for national tender portals, where reading the source in its original language is part of the job.
* Keep them alive: portals change their HTML, move endpoints, break pagination, throttle you.
You find out before the customer does.
* Turn messy sources into clean records: broken HTML, inconsistent XML, APIs that lie about their own schema.
* Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did todays run actually land?" * Work on merge and dedup - the same tender arrives three times, in three shapes, and only one version can reach the customer.
* Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments.
* Guard data quality with tests and checks that fail loudly before a customer finds the gap.
Idóneo profile * Python that holds up: typed, tested, and readable six months later.
* Youve scraped something real: HTTP, HTML and XML parsing, pagination, sessions, rate limits - and you know why a scraper that worked yesterday is broken this morning.




* SQL youre comfortable in: joins, aggregations, window functions.
Youll read from the database every day; tuning and running it isnt your job.
* Builder by default: you see a manual process and your first instinct is to automate it.
* Comfortable with messy sources: broken HTML, inconsistent XML, PDFs that were scans of scans.
* You close your own loop: you check that what you shipped actually ran, before someone else has to ask.
Experience * 1-2 years writing Python in production: scrapers, ETL scripts, automation - anything that had to run unattended and be fixed when it didnt.
* Exposure to an orchestrator (Prefect, Airflow, Dagster) is a plus, not a requirement: youll learn ours properly.
* Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.
What you wont find here * No micromanagement: we trust you to own your part of the stack.
* No "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile.
* No "weve always done it this way" excuses: were here to disrupt, not to follow old patterns.
Our Tech Stack * Data & Infra: Python, PostgreSQL, Prefect, AWS * AI: batch LLM extraction, embeddings, OCR Compensation RAL €35,000 - €45,000 + equity, depending on profile.
Hiring Manager Lorenzo Rossetto ******

📌 Junior Data Engineer - España (Madrid)
🏢 Cato
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: junior data engineer - españa (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: junior data engineer - españa (madrid) / madrid