17 sep
|
Roche
|
San Cugat del Vallés
17 sep
Roche
San Cugat del Vallés
Experteer Overview
As a Site Reliability Engineer at Roche, you will help design and scale reliable distributed systems powering healthcare innovation. You’ll collaborate with development teams to raise reliability, performance, and efficiency across AWS and Azure environments. You lead incident management, postmortems, and continuous improvement to reduce toil and improve uptime. This role offers impact at scale, shaping the backbone of Roche’s IT infrastructure to support patients worldwide. Join a mission-driven team that values diversity and proactive ownership in a fast-paced, on-call environment.
Compensaciones / Incentivos
• Define and implement SLIs, SLOs, and error budgets with product and engineering teams
• Conduct reliability reviews for new and existing services
• Design scalable, fault-tolerant architectures in AWS and Azure
• Lead capacity planning, performance and cost optimization initiatives
• Improve system resilience through automation and self-healing patterns
• Drive observability maturity (metrics, logs, traces, alert quality)
• Perform complex root cause analysis and drive rapid mitigation
• Participate in blameless postmortems and follow-through
• Improve MTTR and production standards
• Collaborate with engineering teams for timely resolutions
• Handle requests and incidents, create and maintain runbooks
• Participate in a structured 24x7 on-call rotation
• Reduce operational toil through tooling and automation (Python or similar)
• Improve CI/CD reliability and deployment safety mechanisms
• Build and maintain infrastructure-as-code (Terraform or equivalent)
• Enhance Kubernetes platform reliability (EKS, AKS, or similar)
• Partner with business, engineering, security, and cloud teams to embed reliability early in SDLC
• Mentor mid-level engineers and shape SRE best practices
• Champion a culture of ownership and continuous improvement
Responsabilidades
• Bachelor in Computer Science or Engineering or equivalent
• Experience in site reliability engineering or software engineering with on-call experience
• Solid experience with AWS and/or Azure including Kubernetes (EKS/AKS/GKE)
• Proficiency with observability tools
• Hands-on incident management tools experience
• Scripting for automation (e.g., Python)
• Proven troubleshooting in cloud and distributed systems
• Excellent communication, teamwork, and documentation skills
Requisitos principales
•
📌 Senior Site Reliability Engineer (San Cugat del Vallés)
🏢 Roche
📍 San Cugat del Vallés