01 ago
|
Allianz
|
Madrid
Experteer Overview
In this Site Reliability Engineer role, you will own the reliability of the central engineering platform within the Advanced Analytics domain. You’ll partner with platform, security, and incident-response teams to meet reliability commitments across AI services, Java APIs, and frontend workloads. You drive automation to eliminate toil, define SLOs/SLIs, and guide incident response and post-incident reviews. The role blends platform engineering, cloud infrastructure, and observability to enable safe, scalable, and cost-conscious delivery.
Compensaciones / Beneficios
• Define, instrument, and maintain SLOs/SLIs for platform components with error-budget tracking and leadership reporting
• Lead on-call escalation and incident response for cluster, network, and storage failures; chair blameless post-incident reviews
• Operate Kubernetes infrastructure (AKS): cluster lifecycle, networking, quotas, autoscaling, and multi-tenancy
• Develop Infrastructure as Code (Terraform) to provision and manage Azure resources with rollbacks
• Build and maintain observability stack: Prometheus, Grafana, Azure Monitor, Application Insights; manage alerts, dashboards, and tracing
• Perform capacity planning and cost-aware resource management across namespaces
• Identify and automate toil through scripting and tooling; measure toil reduction over time
• Maintain platform reliability procedures: upgrades, backup/recovery tests, DR runbooks,
change freeze coordination
• Contribute to CI/CD and GitOps (GitHub Actions, ArgoCD) with reliability-focused release gates and rollback mechanisms
• Collaborate with incident-response and security teams on SRE targets, hardening, and vulnerability remediation
Responsabilidades
• 5+ years in SRE, DevOps, or platform engineering
• Strong Kubernetes experience including cluster operations, networking, storage, and troubleshooting
• Terraform experience; Bicep or ARM templates a plus
• Production Azure expertise (AKS, ACR, Key Vault, Azure Monitor, Application Insights, Virtual Networks, Private Endpoints)
• Observability fundamentals: Prometheus, Grafana, centralized logging, alerting, distributed tracing
• SLO/SLI methodology knowledge; capacity planning experience
• Structured incident management: on-call ownership, blameless post-incident reviews, runbooks
• Scripting/automation in Python or Bash
• CI/CD experience with GitHub Actions and ArgoCD or similar GitOps tooling
• Ability to communicate reliability metrics to non-technical stakeholders
Requisitos principales
• courses and targeted development programs
• global environment with mobility and career progression
• Work Well programs supporting health and wellbeing
• versátil work-life balance
• inclusive, integrity-driven culture
• empowerment and personal development
📌 Site Reliability Engineer (m/f/d) (Madrid)
🏢 Allianz
📍 Madrid