Staff Site Reliability Engineer (AI Platform) (Barcelona)

Staff Site Reliability Engineer (AI Platform) (Barcelona)

04 ago
|
Manychat
|
Barcelona

04 ago

Manychat

Barcelona

ph3Who We Are /h3pCreating content that resonates is great — turning that attention into growth is even better. That's what Manychat does. /ppOur AI-powered automations help creators and brands engage with audiences across Instagram, Messenger, WhatsApp, and TikTok — at the right moment, with the right message, minus the manual work. /ppWe're 400+ people across three continents behind the leading platform for conversational growth, with AI skills shaping how we build, how we work, and how we grow. /ph3Who We're Looking For /h3pWant to shape an AI platform's architecture from day one, instead of just maintaining what someone else built? /ppManychat runs AI features for businesses worldwide, and we're hiring an SRE to own the reliability, performance, and cost of that AI infrastructure, and to raise the bar for how our whole engineering org builds on LLMs. /ppThis isn't a classical SRE role with a bit of AI sprinkled on top. We need someone AI-native: you already understand how modern LLM systems behave in production, things like token throughput, provider rate limits, degraded model quality, and inference latency tails, and you treat all of that as a first-class reliability concern, not an afterthought. /ph3Why This Role Is Worth Your Time /h3ulliThe AI Platform is still young, so you'll shape its architecture, standards, and roadmap from the ground up. /liliThe stakes are real. AI features sit right in the critical path of customer-facing automation, so your work actually matters to the business, not just to a dashboard. /liliYou'll partner directly with the Head of Infrastructure, with real autonomy and visibility into how your decisions play out. /li /ulpSound like the kind of ownership you've been looking for? /ph3What You'll Do /h3ulliOwn reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers. /liliDesign and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails. /liliBuild observability for AI systems:



latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals. /liliDrive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix. /liliRun capacity planning and incident response for inference services; write and improve runbooks and postmortems. /liliScale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features. /li /ulh3TO SHINE IN THIS ROLE /h3h3You’ll Need /h3ulli5+ years in SRE / platform / infrastructure engineering, including production ownership at significant scale. /liliHands‑on experience operating LLM‑backed systems in production: provider APIs (Bedrock, OpenAI, Anthropic, or similar), inference pipelines, self‑hosted or managed model serving. /liliDeep cloud‑native background: AWS, Kubernetes, Terraform/IaC, CI/CD. /liliStrong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non-deterministic systems. /liliProven cost‑optimization work: you can show where you cut cloud or inference spend and how you made cost visible. /liliStaff‑level influence: you've set technical direction beyond your own team and brought others along /li /ulh3It Would Be Great If You Have /h3ulliExperience building or operating an LLM gateway/proxy (e.g., LiteLLM, Kong AI Gateway, custom). /liliExperience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton). /liliExperience with eval pipelines and quality monitoring for LLM outputs. /li /ulh3Why this role /h3ulliGreen‑field ownership: the AI Platform is young; you'll shape its architecture, standards,



and roadmap. /liliReal scale and real stakes: AI features sit in the critical path of customer‑facing automation. /liliDirect partnership with the Head of Infrastructure; high autonomy and visibility. /li /ulh3What We Offer /h3h3We care deeply about your growth, well‑being, and comfort /h3ulliHybrid onboarding to start work remotely and relocation support for you and your family. /liliComprehensive health insurance for both you and your family. /liliProfessional development budget for conference tickets, online courses, and other relevant resources to help you grow. /liliFlexible benefits package: no one‑size‑fits‑all perks here. You get a budget and you decide where it goes, from health and wellbeing to family and setting up your home office. /liliHybrid work and generous, versátil time off — planned with your team, not rationed by a rigid quota. /liliIn‑office perks, including free meals and snacks. /liliCompany‑funded sport activities, annual offsites and team‑building events. /liliAI isn't a perk here — it's how we work. Claude, OpenAI, and more are on by default from day one, and we back teams in adopting whatever makes them faster. /li /ulpManychat is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance. /ppThis commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you’re set up for success. /p /p #J-18808-Ljbffr

📌 Staff Site Reliability Engineer (AI Platform) (Barcelona)
🏢 Manychat
📍 Barcelona

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff site reliability engineer (ai platform) (barcelona) / barcelona

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff site reliability engineer (ai platform) (barcelona) / barcelona