At Roche you can show up as yourself, embraced for the unique qualities you bring. The PositionThe PositionWe are building a global Site Reliability Engineering (SRE) team to support critical commercial and internal platforms and applications. You will influence system design, define reliability standards and reduce operational toil through engineering solutions.We commit ourselves to scientific rigor, unassailable ethics and access to medical innovations for all.
As a seasoned Site Reliability Engineer (SRE) at Roche, you will leverage your deep software engineering expertise to propel our products to new heights of robustness, scalability and reliability. Your MissionDesign and maintain cutting-edge tools, scripts and frameworks that automate repetitive tasks, streamline software deployment and manage expansive systems with unparalleled efficiency. Partner closely with forward-thinking development teams to architect and implement high-performance solutions that elevate system efficiency, optimize resource utilization and enhance deployment processes for superior uptime and user satisfaction.Detect system anomalies, troubleshoot swiftly and conduct thorough root cause analyses to prevent recurring issues.Champion continuous improvement by refining monitoring and alerting mechanisms, conducting insightful post-incident reviews and embedding best practices in software lifecycle management.
Your strategic foresight and meticulous planning will ensure our systems are not only reliable but also superlatively performant.Your Core ResponsibilitiesReliability Engineering & ArchitectureDefine and implement SLIs, SLOs, and error budgets with product and engineering teamsConduct reliability reviews for new and existing servicesDesign scalable, fault-tolerant architectures in AWS and Azure environmentsLead capacity planning, performance and cost optimization initiativesImprove system resilience through automation and self-healing patternsDrive organizational observability maturity (metrics, logs, traces, alert quality)Incident Management & Continuous ImprovementPerform complex root cause analysis and drive rapid mitigationParticipate in blameless postmortems and follow-throughImprove MTTR, reduce incident frequency, and elevate production standardsCollaborate seamlessly with engineering teams to enable timely and effective resolutionsHandle requests and incidents, create and maintain runbooksParticipation in a structured 24*7 on-call rotationAutomation & Platform EngineeringReduce operational toil through tooling and automation (Python or similar)Improve CI/CD reliability and deployment safety mechanismsBuild and maintain infrastructure-as-code (Terraform or equivalent)Enhance Kubernetes platform reliability (EKS, AKS,
or similar)Cross-Functional LeadershipPartner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycleMentor mid-level engineers and help shape SRE best practicesChampioning a culture of ownership, accountability, and continuous improvementWho You Are:Minimum bachelor’s degree in computer science, Engineering, or a related field, or equivalent professional experience.Experience in either site reliability engineering, software engineering or related fields with production on-call experience.Solid experience with AWS and/or Azure, including setting up, monitoring, and maintaining cloud resources (incl. Kubernetes, EKS, AKS, GKE, etc knowledge).Proficiency with observability toolsHands-on experience with incident management toolsProficiency in scripting languages for automation purposesDemonstrated proficiency in troubleshooting, especially in cloud and distributed system environmentsExcellent communication, teamwork and documentation skills, with a proactive and self-motivated approach to improving system reliability and operational efficiencies.Excelling in both spoken and written English communication.Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a general impact.PolandType: Full time
📌 Senior Site Reliability Engineer - DevOps (Remote) (España)
🏢 Roche
📍 España