16 sep
|
Infoser Technological Solutions, Hardware, Big Data Experts
|
Madrid
16 sep
Infoser Technological Solutions, Hardware, Big Data Experts
Madrid
We are looking for an experienced Data Engineer to join an international technology project focused on building and evolving a data-processing platform that handles structured and unstructured information from multiple sources.
You will work on multi-format data ingestion, data normalization, entity reconciliation, relational data modeling and classification workflows , helping transform heterogeneous information into reliable and standardized datasets.
? What will you work on?
- Build and maintain Python pipelines for processing Excel, PDF and XML sources.
- Extract structured records and specification lists from Excel using pandas and openpyxl .
- Process unstructured requirements from PDF documents using pypdf .
- Parse XML documents securely using defusedxml , including protection against XXE vulnerabilities.
- Map raw source fields into a canonical data model .
- Normalize categorical and hierarchical attributes such as product line, variant, region and language.
- Reconcile inconsistent terminology and entity names using fuzzy and approximate matching .
- Maintain relational data models using PostgreSQL and SQLAlchemy ORM .
- Support and improve existing scikit-learn classification workflows .
- Retrain and evaluate models and adjust decision thresholds according to precision, recall and F1 requirements .
?️ What are we looking for?
- Strong professional experience with Python .
- Experience with Django for backend development and data-driven applications.
- Experience working with pandas and openpyxl .
- Experience processing or extracting information from PDF and XML documents .
- Knowledge of data normalization, schema mapping and canonical data models .
- Experience with fuzzy matching techniques, ideally using rapidfuzz .
- Strong knowledge of SQL and PostgreSQL .
- Experience with SQLAlchemy ORM .
- Practical experience with scikit-learn classification workflows .
- Understanding of classification metrics such as precision, recall and F1-score .
- Experience with algorithms such as KNN, Random Forest or SVM .
➕ Nice to have
- Experience with AWS Aurora or other cloud-managed relational databases.
- Experience designing data ingestion or ETL/ELT pipelines.
- Experience working with heterogeneous datasets and complex source-data reconciliation.
- Knowledge of model threshold tuning and classification performance optimization.
? Technology Stack Python · pandas · openpyxl · pypdf · defusedxml · rapidfuzz · PostgreSQL · SQLAlchemy · scikit-learn · KNN · Random Forest · SVM · AWS Aurora
If you enjoy solving complex data-processing problems and working at the intersection of Data Engineering, Data Quality and Machine Learning , we would love to hear from you.
? Apply or contact us to learn more about the project.
📌 Data Engineer - Python & ML (Madrid)
🏢 Infoser Technological Solutions, Hardware, Big Data Experts
📍 Madrid