Senior Site Reliability Engineer (d/f/m) (Madrid)

Senior Site Reliability Engineer (d/f/m) (Madrid)

04 ago
|
Thyssenkrupp Elevator
|
Madrid

04 ago

Thyssenkrupp Elevator

Madrid

ppTK Elevator (TKE) is a global leader in vertical transportation and urban mobility. We provide engineering that keeps the world moving, including design, installation, and maintenance of elevators, escalators, walkways, lifts, passenger boarding bridges, stairlifts, platform lifts and home elevators – including multi-brand modernization and service any place, any time. With TK Elevator’s AI and digital solutions there are no longer any limits to urban mobility. TK Elevator became independent following its separation from the thyssenkrupp group in 2020. The company achieved sales of €9.2 billion in fiscal year 2024/2025. With around 50,000 employees, 25,000 service technicians and over 1,000 support centers globally, we are moved by what moves people. TKE – Move Beyond. /p pTo strengthen our global Digital Technology organization, we are looking for an experienced bSenior Site Reliability Engineer (d/f/m) /bwho will take ownership of System Health Monitoring across our digital ecosystem and the MAX IoT Platform. In this highly visible role, you will help establish reliability as a core capability across the organization by defining observability standards, driving monitoring excellence, and enabling engineering teams to build resilient, reliable solutions. /p h3Drive Monitoring Observability Excellence /h3 ul liOwn and continuously evolve System Health Monitoring across products and platforms. /li liDefine and govern observability standards, monitoring requirements, health models, and alerting strategies. /li liEstablish a unified view of platform health utilizing Azure observability solutions, Central Log Analytics, Grafana, and DevOps monitoring tools. /li liDesign, implement, and optimize Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budget frameworks.



/li /ul h3Enhance Platform Reliability /h3 ul liPromote reliability engineering best practices, including health checks, error-budget management, and post-incident learning processes. /li liAnalyze incident trends and identify opportunities to improve monitoring, alerting, synthetic testing, and operational resilience. /li liTransform monitoring data into actionable insights and predictive analytics that proactively identify risks and performance issues. /li /ul h3Enable Teams Foster Collaboration /h3 ul liPartner closely with Product, Architecture, DevOps, and Incident Operations teams to improve monitoring coverage and alert quality. /li liProvide technical guidance, coaching, and mentorship to engineering teams across the organization. /li liDevelop and maintain high-quality documentation, including runbooks, monitoring standards, and incident-response procedures. /li liContribute actively to the integral DevOps community by sharing knowledge, best practices, and lessons learned. /li /ul ul liMinimum 5 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or a related discipline. /li liStrong expertise in Microsoft Azure and cloud-native technologies. /li liDeep knowledge of: /li ul liAzure Monitor /li liLog Analytics / Kusto Query Language (KQL) /li liApplication Insights /li liGrafana /li /ul liProven experience defining and managing SLIs, SLOs, error budgets,



and reliability frameworks for large-scale distributed systems. /li liStrong understanding of distributed architectures, cloud platforms, and Azure PaaS services. /liliExperience with incident management, post-mortem processes, and continuous reliability improvement. /li liSolid scripting and automation skills using technologies such as C#/.NET, PowerShell, or Python. /li liExcellent analytical, problem-solving, and communication skills. /li liFluency in English (written and spoken) is required. /li /ul h3Nice to Have /h3 ul liExperience with Power BI, Databricks, or AI-driven observability and analytics solutions. /li liKnowledge of Azure DevOps, CI/CD pipelines, Git, and Agile methodologies. /li liBachelor's degree in Computer Engineering or a related technical discipline. /li /ul ul libHealth and Safety /b – Highest standards and a wide range of health promotion and healthcare activities /li libFlexibility /b – We support, for example, through flexible yet regulated working hours and remote working options /li libCollaboration diversity /b – Collegiality is of huge importance – we treat everyone with respect and appreciation /li libDevelopment /b – Individual support to help you get started in your new job as well as training and education programs to help you develop professionally and personally /li libCreative leeway /b – We offer an environment in which you can try out new solutions in a no-blame-culture /li libSustainability /b – We act with responsibility and environmental awareness /li libWork environment /b – We have modern workplaces and IT equipment, subsidized lunchtime meals in the canteen, free parking and discounted public transport tickets /li /ul /p #J-18808-Ljbffr

📌 Senior Site Reliability Engineer (d/f/m) (Madrid)
🏢 Thyssenkrupp Elevator
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer (d/f/m) (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer (d/f/m) (madrid) / madrid