05 ago
|
Allianz Partners
|
Madrid
05 ago
Allianz Partners
Madrid
ph3Key Responsibilities /h3 pAs a Site Reliability Engineer within Advanced Analytics (DA3) in the Chief Data AI Office at Allianz Partners, you will join the platform engineering team to own the reliability and operational health of the central engineering platform. /p pYou will define and maintain service level objectives, drive incident response at the infrastructure layer, and systematically eliminate operational toil through automation. /p pYou will work closely with Platform Engineers, Security Engineers, and incident-response leads to ensure the platform meets its reliability commitments across production workloads spanning AI services, Java APIs, and frontend applications. /p pThrough this role, you will have the main following responsibilities: /p ul liDefine, instrument, and maintain SLOs and SLIs for platform components; own error budget tracking and produce regular reliability reports for senior leadership. /li liServe on the on-call rotation as the infrastructure escalation tier; lead incident response for cluster-level, network-level, and storage failures; chair blameless post-incident reviews. /li liImplement and operate Kubernetes infrastructure (AKS): cluster lifecycle management, networking, resource quotas, autoscaling configuration, and multi-tenancy patterns across product team namespaces. /li liDevelop Infrastructure as Code (Terraform) to provision and manage Azure resources with consistency, auditability, and repeatable rollback capability. /li liBuild and maintain observability infrastructure: Prometheus, Grafana, Azure Monitor, and Application Insights; own alerting rules, dashboards, and distributed tracing coverage across platform components. /li liPerform capacity planning and cost-aware resource management: right-size node pools, tune vertical and horizontal pod autoscalers, and identify resource waste across namespaces. /li liIdentify and eliminate toil: automate repetitive operational tasks through scripting and tooling; measure and track toil reduction over time. /li liMaintain platform reliability procedures: rolling upgrades, backup and recovery testing, disaster recovery runbooks, and change freeze coordination.
/li liContribute to CI/CD pipelines and GitOps tooling (GitHub Actions, ArgoCD) from a reliability and deployment safety perspective; work with platform engineering on release gates and rollback mechanisms. /li liCollaborate with incident-response leads on incident SLA targets and operational procedures; work with Security Engineers on infrastructure hardening and vulnerability remediation. /li /ul h3What You Bring /h3 ul li5+ years professional experience in site reliability engineering, DevOps, or platform engineering roles. /li liStrong Kubernetes experience: cluster operations, networking (Ingress, network policies), storage, autoscaling, and hands‑on troubleshooting across production environments. /li liSolid Infrastructure as Code experience with Terraform; familiarity with Bicep or ARM templates is a plus. /li liProduction experience with Azure cloud services: AKS, ACR, Key Vault, Azure Monitor, Application Insights, Virtual Networks, and Private Endpoints. /li liStrong observability experience: Prometheus, Grafana, centralized logging, alerting configuration, and distributed tracing instrumentation. /li liWorking knowledge of SLO/SLI methodology: error‑budget principles, reliability target setting, and capacity planning. /li liStructured incident management experience: on‑call ownership, blameless post‑incident review, and runbook authorship. /li liScripting and automation proficiency in Python or bash for toil elimination and operational tooling. /li liStrong CI/CD experience: GitHub Actions and ArgoCD or equivalent GitOps tooling. /li /ul h3Ways of Working /h3 ul liComfortable in agile, iterative delivery environments with personal ownership and accountability for platform reliability. /li liClear communicator across global,
cross‑functional stakeholders; able to translate technical reliability metrics into business impact for non‑technical audiences. /li liProactive learner with pragmatic adoption of AI‑assisted developer tools (e.g., GitHub Copilot, Claude Code) to improve automation coverage and delivery velocity. /li /ul h3Nice to Have /h3 ul liKubernetes certifications: CKA or CKAD. /li liExperience supporting AI or ML infrastructure workloads: GPU scheduling, model serving platforms, or inference pipeline operations. /li liExposure to chaos engineering practices and fault injection testing. /li liFinOps experience: reserved capacity planning, resource right‑sizing programs, and cost attribution per team or workload. /li liService mesh experience (Istio, Linkerd) for traffic management and reliability patterns. /li liExperience in regulated industries (insurance, finance, healthcare) where auditability, change traceability, and secure‑by‑default operations are standard practice. /li /ul h3What We Offer /h3 pOur employees play an integral part in our success as a business. We appreciate that each of our employees are unique and have unique needs, ambitions and we enjoy being a part of their journey. We are there to empower and encourage you with your personal and professional development ensuring that you take control by offering a large variety of courses and targeted development programs. /p pAll that in a integral environment where international mobility and career progression are encouraged. Caring for your health and wellbeing is key priority for us. This is why we build Work Well programs to providing you with peace of mind and give the flexibility in planning and arranging for a better work‑life balance. /p p90377 | Data AI | Professional | Allianz Partners | Full‑Time | Permanent /p pWe therefore welcome applications regardless of ethnicity or cultural background, age, gender, nationality, religion, social class, disability or sexual orientation, or any other characteristics protected under applicable local laws and regulations. /p /p #J-18808-Ljbffr
📌 Site Reliability Engineer (m/f/d) (Madrid)
🏢 Allianz Partners
📍 Madrid