03 ago
|
Manychat
|
Barcelona
03 ago
Manychat
Barcelona
Who you are
- 5+ years in SRE / platform / infrastructure engineering, including production ownership at significant scale
- Hands-on experience operating LLM-backed systems in production: provider APIs (Bedrock, OpenAI, Anthropic, or similar), inference pipelines, self-hosted or managed model serving
- Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD
- Strong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non-deterministic systems
- Proven cost-optimization work: you can show where you cut cloud or inference spend and how you made cost visible
- Staff-level influence: you've set technical direction beyond your own team and brought others along
- Experience building or operating an LLM gateway/proxy (e.g., LiteLLM, Kong AI Gateway, custom)
- Experience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton)
- Experience with eval pipelines and quality monitoring for LLM outputs
What the job involves
- Want to shape an AI platform's architecture from day one, instead of just maintaining what someone else built?
- Manychat runs AI features for businesses worldwide, and we're hiring an SRE to own the reliability, performance, and cost of that AI infrastructure, and to raise the bar for how our whole engineering org builds on LLMs
- This isn't a classical SRE role with a bit of AI sprinkled on top. We need someone AI-native: you already understand how modern LLM systems behave in production, things like token throughput, provider rate limits, degraded model quality, and inference latency tails, and you treat all of that as a first-class reliability concern, not an afterthought
- The AI Platform is still young, so you'll shape its architecture, standards, and roadmap from the ground up
- The stakes are real. AI features sit right in the critical path of customer-facing automation, so your work actually matters to the business, not just to a dashboard
- You'll partner directly with the Head of Infrastructure, with real autonomy and visibility into how your decisions play out
- Sound like the kind of ownership you've been looking for?
- Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers
- Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails
- Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals
- Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix
- Run capacity planning and incident response for inference services; write and improve runbooks and postmortems
- Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features
Benefits
- An opportunity for rapid personal and professional growth
- Uncapped Commission + Equity
- Generous time off policy to balance your work and life, including paid parental leave
- Competitive medical, dental, and vision coverage for you and your dependents
- Annual wellness reimbursement
- Collaborative, transparent, and fun-loving office culture
- We pay for relevant conference tickets, training programs, courses, and any necessary literature
📌 Staff Site Reliability Engineer (Barcelona)
🏢 Manychat
📍 Barcelona