Staff Incident Manager (Madrid)

Staff Incident Manager (Madrid)

04 ago
|
Yuno
|
Madrid

04 ago

Yuno

Madrid

ppUS · Europe · APAC · Remote · Full Time · Individual Contributor · +7 Years of Experience /p h3Who We Are /h3 pAt Yuno, we are building the payment infrastructure that allows all companies to participate in the global market. Founded by seasoned experts from the payments and tech industries, our technology provides access to leading payment capabilities, enabling companies to engage customers confidently and maintain global operations through seamless integrations. /p pWe empower high-performing teams at brands like InDrive, McDonald’s, Rappi, and Viva Aerobus to integrate over 1,000 payment methods via a single API. By leveraging advanced AI and the latest technologies, we orchestrate smart routing and fraud prevention across 80+ countries. /p h3About The Role /h3 pWe are hiring a Staff Incident Manager to own the major incident lifecycle end to end. You are the person who takes control when something is broken in production, brings the right people together, drives the response to resolution, and makes sure we are measurably better after every incident than we were before it. /p pThis is not a passive coordination role. You set the standard for how Yuno detects, responds to, communicates about, and learns from incidents. You will spend as much time fixing the process as you do running the response. /p pOn-call responsibility is core to this role. Yuno's engineering teams run a “You Build It, You Run It” model – pods own their own on-call. The Incident Commander function exists to coordinate cross-domain and Sev-1 events where multiple pods are involved. /p h3Your Contribution Will Be /h3 ul lipbIncident command — /bact as incident commander on major and critical incidents, owning coordination, decision‑making cadence, and escalation from detection through resolution; treat merchant transaction impact, PSP and acquirer dependency failures, settlement and reconciliation knock‑on effects, and PCI‑DSS scope as first‑class concerns in every response. /p /li lipbReliability metrics — /bdrive down time to detect, time to engage, and time to recover; own MTTR as a headline metric and the operational discipline behind Yuno's 99.99% uptime target. /p /li lipbOn‑call program — /brun rotations, escalation policies, paging hygiene, and alert quality; reduce noise so on‑call engineers trust their pages. /p /li lipbIncident communications — /bown internal stakeholder alignment during an event and drive clear,



accurate merchant‑facing updates — including the status page — in coordination with Support and account teams. /p /li lipbPostmortems — /brun blameless postmortems, hold the room to a no‑blame standard, and make sure action items are concrete, owned, and tracked to closure. /p /li lipbReliability roadmap — /btranslate recurring incident patterns into reliability work, partnering with engineering teams to turn postmortem findings into roadmap items, not orphaned tickets. /p /li lipbOperating model — /bdefine and maintain incident severity levels, response runbooks, and the operating model for declaring and managing incidents. /p /li lipbReporting — /breport on incident trends, reliability posture, and SLA and SLO performance to engineering leadership. /p /li /ul h3Skills You Need /h3 h3Minimum Qualifications /h3 ul lipProven experience running major incident response in a production environment that real customers depend on, ideally as an incident commander or in a dedicated incident management function. /p /li lipStrong working knowledge of modern distributed systems and cloud‑native operations — comfortable holding your own on a bridge call with senior engineers during an outage, following the technical thread, and keeping the response moving without needing every detail spoon‑fed. /p /li lipExperience defining or maturing an incident management practice from the ground up: severity frameworks, operating models, runbooks, and on‑call programs built to last, not just inherited. /p /li lipFluency with observability and incident tooling — metrics, logging, tracing, alerting, and paging platforms. Hands‑on experience with Datadog is strongly preferred; familiarity with a paging platform (PagerDuty, OpsGenie, or similar) and a status page tool is expected. /p /li lipA track record of running postmortems that change behavior, and of closing the loop between incidents and engineering work. /p /li lipExcellent written and verbal communication — able to write a clear merchant‑facing status update and a precise internal escalation under pressure. /p /li lipCalm,



decisive judgment during high‑pressure events — you hold structure when others are reacting. /p /li lipComfort being on‑call as a regular part of the role, including for critical incidents outside business hours. /p /li lipEnglish — advanced proficiency required. /p /li /ul h3Preferred Qualifications /h3 ul lipFamiliarity with PCI‑DSS and the operational obligations that come with handling payment flows at scale. /p /li lipSpanish — business‑level proficiency preferred; Yuno's engineering organization spans LatAm, and on‑call bridges and postmortem discussions often run in Spanish. /p /li /ul h3Nice to Have /h3 ul lipPrior experience in payments, fintech, or another high‑availability, regulated domain. /p /li lipExposure to SRE practices, error budgets, and SLO‑driven prioritization. /p /li lipExperience managing incidents across multi‑timezone, remote‑first engineering organizations. /p /li /ul h3What Success Looks Like /h3 ul lipbFirst 90 days — /byou have mapped how incidents currently get declared, run, and reviewed, identified the biggest gaps, and started closing them. Severity definitions and the incident operating model are clear and adopted. /p /li lipbSix months — /bmajor incidents run to a consistent standard. Postmortem action items are tracked and closed. On‑call noise is down and trust in paging is up. Leadership has a reliable view of reliability trends. /p /li lipbOne year — /bthe incident practice is self‑sustaining. Runbooks are owned by engineering pods, not just by you. MTTR and repeat‑incident rates are trending measurably down, and reliability work sourced from incident data is landing on engineering roadmaps and shipping. /p /li /ul h3What We Offer at Yuno /h3 ul liCompetitive Compensation /li liRemote Work - You can work from everywhere! /li liHome Office Bonus - A one‑time allowance to help you create your adecuado home office /li liWork Equipment /li liStock Options /li liHealth Plan wherever you are /li liFlexible Days Off /li liLanguage, Professional, and Personal Growth courses /li /ul pWe may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed or wish to exercise your data protection rights, please contact us at /p /p #J-18808-Ljbffr

📌 Staff Incident Manager (Madrid)
🏢 Yuno
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff incident manager (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: staff incident manager (madrid) / madrid