14 ago
|
Printify
|
Barcelona
14 ago
Printify
Barcelona
Experteer Overview
As Senior SRE II at FYUL Platform Infrastructure, you will own and drive reliability and scalability across multi-account AWS, Kubernetes (EKS), and core platform services. You will design large-scale automation, set engineering standards, and mentor junior SREs, while balancing hands-on platform work with technical leadership. You’ll lead cost optimization, security, and observability improvements to enable product teams to ship reliably at integral scale. This role offers meaningful impact through DevOps enablement and cross-team collaboration.
Compensaciones / Beneficios
• Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts using infrastructure as code
• Operate and design Amazon EKS clusters with networking, storage, and scaling strategies for containerized workloads
• Own core platform services like cloud networking, Kubernetes, databases, and messaging systems
• Drive large-scale automation with Terraform/Terragrunt and GitOps (ArgoCD); establish standards across teams
• Lead adoption of automation to reduce manual work and promote repeatability
• Lead on-call and incident response; author runbooks, ADRs, and postmortems
• Improve reliability and observability with Grafana/Prometheus/Loki/Tempo/Mimir for scalable systems
• Mentor mid-level SREs and support onboarding; communicate complex concepts to engineers and non-technical stakeholders
• Collaborate with product squads to represent Platform Infrastructure in cross-team initiatives
• Drive security (IAM; encryption; secure logging) and FinOps for cost efficiency; audit spend and optimize resources
• Contribute to platform/DevEx roadmap and self-service initiatives for other teams
Responsabilidades
• Solid Linux systems administration and Python scripting
• Strong AWS knowledge (EKS, IAM, VPC, RDS, S3, SQS); multi-account environments a plus
• Hands-on Kubernetes (EKS) operations, Helm, CNI (Cilium), IPAM, container security (ECR)
• Terraform (modules, state) and Terragrunt; GitOps with ArgoCD
• Postgres/MySQL/MongoDB in production; Aurora experience a plus
• CI/CD with Jenkins and/or GitHub Actions; blue/green and canary deployment experience
• Grafana/Prometheus/Loki/Tempo/Mimir for observability; dashboarding, alerting, tracing
• Incident management experience; on-call rotations, runbooks, postmortems
• Knowledge of 12-Factor App principles and FinOps awareness
• Nice to have: GCP, Kafka/AWS MSK, regulated environments, self-service platform experience
• Languages: PHP, Node.js; Infra: AWS, Kubernetes, Terraform, Helm, Atlantis
Requisitos principales
• remote work options
• private health insurance
• extra days off for wellbeing
• annual learning opportunities
• mentorship and internal meetups
• office lunch in Riga
📌 Senior Site Reliability Engineer (remote within EMEA) (Barcelona)
🏢 Printify
📍 Barcelona