Staff Site Reliability Engineer — AI Platform (Barcelona)

Staff Site Reliability Engineer — AI Platform (Barcelona)

25 ago
|
Manychat
|
Barcelona

25 ago

Manychat

Barcelona

Manychat is a leading chat marketing platform, helping businesses engage with their customers on Instagram, Facebook Messenger, WhatsApp, and Telegram. Trusted by over 1 million brands in 170+ countries, we’re an official Meta Business Partner, backed by Bessemer Venture Partners. Manychat is a leading chat marketing platform, helping businesses engage with their customers on Instagram, Facebook Messenger, WhatsApp, and Telegram. Trusted by over 1 million brands in 170+ countries, we’re an official Meta Business Partner, backed by Bessemer Venture Partners. Manychat runs AI features for businesses worldwide, and we’re hiring an SRE to own the reliability, performance, and cost of that AI infrastructure, and to raise the bar for how our whole engineering org builds on LLMs.
We need someone AI-native: you already understand how modern LLM systems behave in production, things like token throughput, provider rate limits, degraded model quality, and inference latency tails, and you treat all of that as a first-class reliability concern, not an afterthought.

Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers.
Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals.
Run capacity planning and incident response for inference services; 5+ years in SRE / platform / infrastructure engineering,



including production ownership at significant scale.
~ Deep cloud‑native background: AWS, Kubernetes, Terraform/IaC, CI/CD.
~ Strong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non‑deterministic systems.
~ Proven cost‑optimization work: you can show where you cut cloud or inference spend and how you made cost visible.
~ LiteLLM, Kong AI Gateway, custom).
Experience with eval pipelines and quality monitoring for LLM outputs.

Hybrid onboarding to start work remotely and relocation support for you and your family.
Comprehensive health insurance for both you and your family.
Professional development budget for conference tickets, online courses, and other relevant resources to help you grow.
You get a budget and you decide where it goes, from health and wellbeing to family and setting up your home office.
Hybrid work and generous, adaptable time off — planned with your team, not rationed by a rigid quota.
In‑office perks, including free meals and snacks.
We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance.

📌 Staff Site Reliability Engineer — AI Platform (Barcelona)
🏢 Manychat
📍 Barcelona

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff site reliability engineer — ai platform (barcelona) / barcelona

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff site reliability engineer — ai platform (barcelona) / barcelona