Senior SRE EngineerJoin DempoOwn reliability at scale with usWe are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience. This role sits at the Senior level on our Cloud/Platform/SRE career path — reliability engineering with a heavy focus on metrics and production systems. You'll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go.You'll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load.ResponsibilitiesReliability & Incident Management· Lead incidents as commander: set and revise severity, and know when to mitigate first and diagnose later· Own the incident record and timeline standard,
including the link between deployments and incidents· Communicate with stakeholders while an incident is active· Conduct blameless postmortems and drive toil identification and elimination as measured workObservability· Implement the three pillars of observability (logs, metrics, traces) end to end· Design metrics and query strategy — recording rules, dashboard design that separates on-call needs from analyst needs· Define SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one that's easiest to measure· Design alerting systems — routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts)Platform & Production Systems· Design workload health signals — liveness, readiness, and startup probes — and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining)· Build runbook automation and self-healing systems to reduce operational toil· Contribute to CI/CD framework improvements and cost optimization initiativesTechnical Leadership· Make architectural decisions for reliability-critic
📌 Senior Sre Engineer (España)
🏢 Dempo
📍 España