Staff Site Reliability Engineer (AI Platform) (Barcelona)

Staff Site Reliability Engineer (AI Platform) (Barcelona)

04 ago
|
Doist
|
Barcelona

04 ago

Doist

Barcelona

ph3WHO WE ARE /h3pWe help creators get more out of every conversation with Instagram-focused automations and support for other channels like Messenger, WhatsApp, and TikTok. The result? Better engagement, more sales, and real, sustainable growth. With a diverse team of 350+ people spread across three continents, we're building the leading Chat Marketing platform that is used - and loved - by more than 1.5 million customers worldwide. /ph3WHO WE'RE LOOKING FOR /h3pWe're looking for a Senior Site Reliability Engineer who thrives at the crossroads of classic Linux and AWS infrastructure and modern Site Reliability Engineering. This is a high-impact, hybrid role designed for someone who can manage cloud resources, harden Kubernetes clusters, and shape a more reliable and developer-friendly platform. We need you not just to maintain but to rethink and evolve our infrastructure, balancing hands-on operations with strategic improvements that future-proof our growing AI product landscape. You'll take over key responsibilities from our current Infra Lead who is transitioning to a software-focused role, giving you immediate ownership and space to shine. /ph3WHY THE ROLE IS SPECIAL /h3pYou won't be a cog in a massive SRE org. You'll be the bridge between Infrastructure and Engineering, shaping how we scale Kubernetes, how we approach platform reliability, and how developers ship fast without fear. You'll get autonomy, ownership, and a smart, humble team excited to learn with you. /ph3WHAT YOU'LL DO /h3ulliMaintain and harden AWS infrastructure (EC2, ALB/NLB, WAF, IAM, CloudWatch)



/liliOperate and evolve our EKS clusters powering Python-based AI services /liliMigrate existing services to Kubernetes using Terraform and Helm /liliCodify infrastructure with Terraform and manage host-level automation via Ansible /liliBuild and improve CI/CD pipelines with GitHub Actions /liliOwn observability efforts: Prometheus, Grafana, alerting, and on-call readiness /liliSupport OS-level patching, certs, WAF rules, and general infra hygiene /liliPartner with engineers to guide best practices and drive platform reliability /liliCreate clean, maintainable infrastructure documentation and playbooks /liliOccasionally support rare off-hours incidents (don’t worry, really rare) /li /ulh3TO SHINE IN THIS ROLE /h3ulli5+ years of experience managing Linux in production (Ubuntu, Amazon Linux) /liliStrong experience with Kubernetes (ideally EKS), Helm, and Terraform /liliComfort with running and debugging Python workloads in containers /liliSolid understanding of networking, IAM, and cloud security best practices /liliHands-on Nginx experience (Ingress and reverse proxy setups) /liliExcellent communication skills; you can explain complex infra to devs clearly /li /ulh3NICE TO HAVE SKILLS /h3ulliStrong Ansible skills beyond the basics /liliPostgreSQL or Amazon RDS tuning and operations experience /liliDeep understanding of observability tools (Prometheus, Grafana, Loki, etc.) /liliFamiliarity with PHP production environments /liliExperience with TDD,



CI/CD best practices, and agile development /liliAny previous SRE-like exposure such as building resilience, automation, or incident tooling /li /ulh3WHAT WE OFFER /h3ulliHybrid onboarding to start work remotely and relocation support for you and your family. /liliComprehensive health insurance for both you and your family. /liliProfessional development budget for conference tickets, online courses, and other relevant resources to help you grow. /liliFlexible benefits package to tailor perks that matters most for you. /liliHybrid work and generous leave options to prioritize your work-life balance. /liliIn-office perks, including free meals and snacks. /liliCompany-funded sport activities, annual offsites and team-building events. /li /ulpManychat is an Equal Opportunity Employer. We're committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance. This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you're set up for success. /ppWith my application, I accept the Manychat Privacy Policy . /p /p #J-18808-Ljbffr

📌 Staff Site Reliability Engineer (AI Platform) (Barcelona)
🏢 Doist
📍 Barcelona

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff site reliability engineer (ai platform) (barcelona) / barcelona

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff site reliability engineer (ai platform) (barcelona) / barcelona