We are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience. This role sits at the Senior level on our Cloud/Platform/SRE career path — reliability engineering with a heavy focus on metrics and production systems. You’ll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go.You’ll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load.ResponsibilitiesReliability & Incident ManagementLead incidents as commander: set and revise severity, and know when to mitigate first and diagnose laterOwn the incident record and timeline standard, including the link between deployments and incidentsCommunicate with stakeholders while an incident is activeConduct blameless postmortems and drive toil identification and elimination as measured workObservabilityImplement the three pillars of observability (logs, metrics,
traces) end to endDesign metrics and query strategy — recording rules, dashboard design that separates on-call needs from analyst needsDefine SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one that's easiest to measureDesign alerting systems — routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts)Platform & Production SystemsDesign workload health signals — liveness, readiness, and startup probes — and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining)Build runbook automation and self-healing systems to reduce operational toilContribute to CI/CD framework improvements and cost optimization initiativesTechnical LeadershipMake architectural decisions for reliability-critical systemsMentor cloud/platform engineersInfluence technical direction on infrastruct
📌 Senior SRE Engineer (España)
🏢 Dempo
📍 España