03 ago
|
Allianz
|
Madrid
Experteer Overview
¿Es este el siguiente paso en su carrera? Descubra si es el candidato adecuado leyendo la descripción completa a continuación.
In this Site Reliability Engineer role, you will own the reliability of the central engineering platform within the Advanced Analytics domain. You’ll partner with platform, security, and incident-response teams to meet reliability commitments across AI services, Java APIs, and frontend workloads. You drive automation to eliminate toil, define SLOs/SLIs, and guide incident response and post-incident reviews. The role blends platform engineering, cloud infrastructure, and observability to enable safe, scalable, and cost-conscious delivery.
Compensaciones / Incentivos
• Define, instrument, and maintain SLOs/SLIs for platform components with error-budget tracking and leadership reporting
• Lead on-call escalation and incident response for cluster, network, and storage failures; chair blameless post-incident reviews
• Operate Kubernetes infrastructure (AKS): cluster lifecycle, networking, quotas, autoscaling, and multi-tenancy
• Develop Infrastructure as Code (Terraform) to provision and manage Azure resources with rollbacks
• Build and maintain observability stack: Prometheus, Grafana, Azure Monitor, Application Insights; manage alerts, dashboards, and tracing
• Perform capacity planning and cost-aware resource management across namespaces
• Identify and automate toil through scripting and tooling; measure toil reduction over time
• Maintain platform reliability procedures: upgrades, backup/recovery tests, DR runbooks, change freeze coordination
• Contribute to CI/CD and GitOps (GitHub Actions, ArgoCD) with reliability-focused release gates and rollback mechanisms
• Collaborate with incident-response and security teams on SRE targets, hardening, and vulnerability remediation
Responsabilidades
• 5+ years in SRE, DevOps, or platform engineering
• Strong Kubernetes experience including cluster operations, networking, storage, and troubleshooting
• Terraform experience; Bicep or ARM templates a plus
• Production Azure expertise (AKS, ACR, Key Vault, Azure Monitor, Application Insights, Virtual Networks, Private Endpoints)
• Observability fundamentals: Prometheus, Grafana, xqbhyrx centralized logging, alerting, distributed tracing
• SLO/SLI methodology knowledge; capacity planning experience
• Structured incident management: on-call ownership, blameless post-incident reviews, runbooks
• Scripting/automation in Python or Bash
• CI/CD experience with GitHub Actions and ArgoCD or similar GitOps tooling
• Ability to communicate reliability metrics to non-technical stakeholders
Requisitos principales
• courses and targeted development programs
• global environment with mobility and career progression
• Work Well programs supporting health and wellbeing
• flexible work-life balance
• inclusive, integrity-driven culture
• empowerment and personal development
📌 Site Reliability Engineer (m/f/d) (Madrid)
🏢 Allianz
📍 Madrid