05 oct
|
AstraZeneca
|
Barcelona
05 oct
AstraZeneca
Barcelona
Overview
In this role you will lead the design, implementation, and optimization of a cloud-based data lakehouse to empower products, data science, and agents. You will be one of the senior engineers, providing hands-on work and mentoring to the team while aligning data infrastructure with security and product goals. You’ll drive governance, scalability, and reliability across the data foundation, leveraging cutting-edge AWS technologies and modern data tooling. The mission focuses on transforming life sciences data to enable faster, more impactful outcomes.
Compensaciones / Beneficios
in-person collaboration three days per week in Barcelona versátil working arrangement
Responsabilidades
Design and manage AWS data services including Lake Formation, Glue, Athena, and EMR/Redshift Serverless to build an integrated data foundation Work with Open Table Formats (S3 Tables, Apache Iceberg preferred, or Delta Lake) focusing on partition, schema evolution, time travel and compaction Build and operate streaming pipelines with Kinesis Data Streams or MSK, ensuring exactly-once semantics and handling late-arriving data Define infrastructure as code (AWS CDK TS or CloudFormation); enable CI/CD for data pipelines with GitHub Actions and Terraform Model data with dimensional schemas and slowly changing dimensions; balance normalization vs denormalization for access patterns Implement governance and security including column-level and row-level controls and policy-driven data classification Develop ETL logic and data quality checks using Python or Spark; utilize PySpark or Spark Scala for distributed transforms Foster AI/ML exposure and practical tooling adoption; provide mentorship and technical leadership across tooling and automation Collaborate with product management and security to align data strategies with business goals and cohesive workflows Mentor junior and mid-level engineers and promote a learning-oriented, collaborative culture
Requisitos principales
10+ years in Data Engineering, with SaaS and multi-tenant data platforms experience Strong AWS expertise (VPC, IAM, EC2, S3, RDS, Lambda, EKS, WAF, CloudTrail) Expert knowledge of S3, RDS, DynamoDB, Kinesis, Glue, DataZone, Athena, Redshift Serverless, and EventBridge Proficiency with Docker, Kubernetes, Helm; CI/CD with ArgoCD and GitHub Actions; IaC with AWS CDK TS and CloudFormation Security literacy including IAM, KMS, encryption standards, NIST frameworks Experience with OpenTelemetry, Prometheus, Grafana, CloudWatch, and CloudTrail monitoring Extensive knowledge of AWS Glue, Kinesis, and Managed Kafka for real-time and batch data processing Strong programming and automation skills in TypeScript and Bash; PySpark or Spark Scala for data processing Multi-account AWS management experience using AWS Control Tower Excellent communication skills for stakeholder alignment and cross-functional collaboration Mentorship and leadership Strategic thinking and big-picture orientation Collaborative teamwork in cross-functional settings AWS Lake Formation, Glue, Athena, EMR/Redshift Serverless Apache Iceberg or Delta Lake S3 Tables, partition/schema evolution, time travel, compaction
📌 Data Foundation Engineering Lead - Evinova (Barcelona)
🏢 AstraZeneca
📍 Barcelona