02 ago
|
Manychat
|
Barcelona
02 ago
Manychat
Barcelona
Experteer Overview
In this role you will own the reliability, performance, and cost of Manychat's AI infrastructure. You’ll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps. You’ll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability. This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company.
Compensaciones / Beneficios
• Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)
• Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails
• Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals
• Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies
• Conduct capacity planning and incident response; develop runbooks and postmortems
• Scale AI expertise across the org by setting standards and coaching teams on LLM-backed features
Responsabilidades
• 5+ years in SRE / platform / infrastructure engineering with production ownership at scale
• Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)
• Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD
• Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems
• Proven cost-optimization experience and ability to make costs visible
• Staff-level influence and ability to drive technical direction beyond own team
Requisitos principales
• Hybrid onboarding with remote start
• Relocation support for you and family
• Comprehensive health insurance
• Professional development budget
• Adaptable benefits package
• Hybrid work and generous time off
📌 Staff Site Reliability Engineer (AI Platform) (Barcelona)
🏢 Manychat
📍 Barcelona