04 ago
|
Thetaray
|
Madrid
ppLocation: ThetaRay Madrid, Community of Madrid, Spain /p h3About ThetaRay /h3 pThetaRay is a trailblazer in AI‑powered Anti‑Money Laundering (AML) solutions, offering cutting‑edge technology to fintechs, banks, and regulatory bodies worldwide. Our mission is to enhance trust in financial transactions, ensuring compliant and innovative business growth. /p h3Why Join ThetaRay? /h3 pAt ThetaRay, you'll be part of a dynamic general team committed to redefining the financial services sector through technological innovation. You will contribute to creating safer financial environments and have the opportunity to work with some of the brightest minds in AI, ML, and fintech. We offer a collaborative, inclusive, and forward‑thinking work environment where your ideas and contributions are valued and encouraged. /p h3Position: Data Engineer /h3 pAs a Data Engineer you will design, implement, and optimize data pipeline flows within the ThetaRay system, supporting data scientists with data flow implementation based on their feature design and constructing complex rules to detect money‑laundering activity. You will build pipeline solutions from the ground up, support multiple production implementations and train customer data scientists and engineers. /p h3Responsibilities /h3 ul liImplement and maintain production data pipelines in the ThetaRay system based on data‑scientist designs. /li liDesign and implement solution‑based data flows for specific use cases, enabling product applicability. /li liBuild a machine‑learning data pipeline. /li liCreate data tools for analytics and data‑science team members. /li liCollaborate with product, RD, data, and analytics experts to enhance system functionality.
/li liTrain customer data scientists and engineers to maintain and amend pipelines. /li liTravel to customer locations domestically and abroad. /li liBuild and manage technical relationships with customers and partners. /li /ul h3Requirements /h3 ul li2+ years of hands‑on experience with Apache Spark. /li liHands‑on experience with SQL. /li liVersion control experience (Git). /li liExperience with the Apache Hadoop ecosystem (Hive, Impala, Hue, HDFS, Sqoop, etc.). /li liPython (Pandas) experience. /li liExperience with PySpark/Scala/Java/R. /li liHands‑on experience with data transformation, validation, cleansing, and ML feature engineering. /li liBSc or higher in Computer Science, Statistics, Informatics, Information Systems, Engineering, or another quantitative field. /li liExperience optimizing big‑data pipelines, architectures, and data sets (advantage). /li liStrong analytical skills with structured and semi‑structured datasets. /li liExperience building processes for data transformation, structures, metadata, dependency, and workload management. /li liRoot‑cause analysis experience on internal and external data and processes. /li liBusiness‑oriented, able to work with external customers and cross‑functional teams. /li liFluent in English and Spanish (written spoken). /li /ul h3Nice to Have /h3 ul liLinux proficiency. /li liExperience building machine‑learning pipelines. /li liExperience with Elasticsearch. /li liExperience with Zeppelin or Jupyter. /liliExperience with workflow automation platforms (Jenkins, Apache Airflow). /li liExperience with microservices architecture, Docker, and Kubernetes. /li /ul h3Employment Details /h3 pSeniority level: Not applicablebr/Employment type: Full‑timebr/Job function: Information Technologybr/Industries: Software Development /p /p #J-18808-Ljbffr
📌 Data Engineer (Madrid)
🏢 Thetaray
📍 Madrid