15 sep
|
Infoser New Technologies
|
Madrid
15 sep
Infoser New Technologies
Madrid
We are looking for an experienced Data Engineer to join an international technology project focused on building and evolving a data-processing platform that handles structured and unstructured information from multiple sources.
You will work on multi-format data ingestion, data normalization, entity reconciliation, relational data modeling and classification workflows, helping transform heterogeneous information into reliable and standardized datasets.
? What will you work on?
- Build and maintain Python pipelines for processing Excel, PDF and XML sources.
- Extract structured records and specification lists from Excel using pandas and openpyxl.
- Process unstructured requirements from PDF documents using pypdf.
- Parse XML documents securely using defusedxml, including protection against XXE vulnerabilities.
- Map raw source fields into a canonical data model.
- Normalize categorical and hierarchical attributes such as product line, variant, region and language.
- Reconcile inconsistent terminology and entity names using fuzzy and approximate matching.
- Maintain relational data models using PostgreSQL and SQLAlchemy ORM.
- Support and improve existing scikit-learn classification workflows.
- Retrain and evaluate models and adjust decision thresholds according to precision, recall and F1 requirements.
?️ What are we looking for?
- Strong professional experience with Python.
- Experience with Django for backend development and data-driven applications.
- Experience working with pandas and openpyxl.
- Experience processing or extracting information from PDF and XML documents.
- Knowledge of data normalization, schema mapping and canonical data models.
- Experience with fuzzy matching techniques, ideally using rapidfuzz.
- Strong knowledge of SQL and PostgreSQL.
- Experience with SQLAlchemy ORM.
- Practical experience with scikit-learn classification workflows.
- Understanding of classification metrics such as precision, recall and F1-score.
- Experience with algorithms such as KNN, Random Forest or SVM.
➕ Nice to have
- Experience with AWS Aurora or other cloud-managed relational databases.
- Experience designing data ingestion or ETL/ELT pipelines.
- Experience working with heterogeneous datasets and complex source-data reconciliation.
- Knowledge of model threshold tuning and classification performance optimization.
? Technology Stack Python · pandas · openpyxl · pypdf · defusedxml · rapidfuzz · PostgreSQL · SQLAlchemy · scikit-learn · KNN · Random Forest · SVM · AWS Aurora
If you enjoy solving complex data-processing problems and working at the intersection of Data Engineering, Data Quality and Machine Learning, we would love to hear from you.
? Apply or contact us to learn more about the project.
📌 Data Engineer - Python & ML (Madrid)
🏢 Infoser New Technologies
📍 Madrid