Experteer Overview In this role you will own and evolve system health monitoring across our digital ecosystem, establishing observability standards and driving reliable platform practices. You will work with cross-functional teams to implement SLIs/SLOs and alerting, leveraging Azure observability tools to deliver proactive insights. Your work supports scaling and resilience for the MAX IoT Platform and related products. This is a high-visibility opportunity to shape reliability culture across a global organization.Compensaciones / Incentivos
- Own and evolve System Health Monitoring across products and platforms
- Define and govern observability standards, monitoring requirements, health models, and alerting strategies
- Unify platform health views using Azure observability solutions, Log Analytics, Grafana, and DevOps monitoring tools
- Design and optimize SLIs, SLOs, and error budgets
- Promote reliability engineering practices including post-incident learning
- Analyze incident trends and improve monitoring, alerting, testing, and resilience
- Translate monitoring data into actionable insights and predictive analytics
- Collaborate with Product, Architecture, DevOps,
and Incident Operations teams on monitoring coverage and alert quality
- Provide guidance and mentorship to engineering teams; maintain runbooks and incident procedures
- Contribute knowledge to the general DevOps communityResponsabilidades
- Min. 5 years in Site Reliability Engineering, DevOps, Cloud Operations, or related field
- Strong expertise in Microsoft Azure and cloud-native tech
- Deep knowledge of Azure Monitor, Log Analytics/KQL, Application Insights, Grafana
- Experience defining/managing SLIs, SLOs, and reliability frameworks for large-scale systems
- Understanding of distributed architectures, cloud platforms, and Azure PaaS services
- Experience with incident management, post-mortems, and reliability improvements
- Scripting/automation skills (C#/.NET, PowerShell, or Python)
- Excellent analytical, problem-solving, and communication skills
- Fluent English (written and spoken)Requisitos principales
- health and safety programs
- flexible working hours
- remote working options
- training and education programs
- modern workplaces and IT equipment
- subsidized meals and discounted transport tickets
📌 Senior Site Reliability Engineer (D/F/M) (Madrid)
🏢 Talent
📍 Madrid