Senior SRE Engineer Join Dempo Own reliability at scale with us We are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience. This role sits at the Senior level on our Cloud/Platform/SRE career path — reliability engineering with a heavy focus on metrics and production systems. Youll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go. Youll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load. Responsibilities Reliability & Incident Management · Lead incidents as commander: set and revise severity, and know when to mitigate first and diagnose later · Own the incident record and timeline standard,
including the link between deployments and incidents · Communicate with stakeholders while an incident is active · Conduct blameless postmortems and drive toil identification and elimination as measured work Observability · Implement the three pillars of observability (logs, metrics, traces) end to end · Design metrics and query strategy — recording rules, dashboard design that separates on-call needs from analyst needs · Define SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one thats easiest to measure · Design alerting systems — routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts) Platform & Production Systems · Design workload health signals — liveness, readiness, and startup probes — and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining) · Build runbook automation and self-healing systems to reduce operational toil · Contribute to CI/CD framework improvements and cost optimization initiatives Technical Leadership · Make architectural decisions for
📌 Senior SRE Engineer (España)
🏢 Dempo
📍 España