29 ago
|
AstraZeneca
|
Barcelona
29 ago
AstraZeneca
Barcelona
Do you have expertise in, and passion for, AI-powered operational excellence? Would you like to apply your expertise to impact a company that follows the science and turns ideas into life changing medicines? If so, AstraZeneca might be the one for you!About AstrazenecaAstraZeneca is a global, innovation-driven BioPharmaceutical business that focuses on the discovery, development and commercialisation of prescription medicines for some of the world’s most serious diseases. But we’re more than one of the world’s leading pharmaceutical companies.At AstraZeneca we’re dedicated to being a Great Place to Work. Where you are empowered to push the boundaries of science and unleash your entrepreneurial spirit. There’s no better place to make a difference to medicine, patients and society. An inclusive culture that champions diversity and collaboration. Always committed to lifelong learning, growth and development.About Our Enterprise Ai Platforms & Services TeamA place to do important work. We connect across the whole business to power each function to better influence patient outcomes and improve their lives. Impactful and valuable, this is where you come to raise your profile and do good for others. Play an increasingly crucial role in driving disruptive transformation on our journey to becoming a digital and data-led enterprise. Unleash the power of our latest innovations in data, machine learning and technology to turn complex information into life-changing and practical insights. Work in synergy with leading experts in our specialist communities. Here we bring the brightest minds to bear, with access to cutting-edge techniques and the opportunity to be part of novel solutions. It’s up to us to drive the outcomes forward, inventing and building, expanding our knowledge to identify the next opportunity. An inclusive team, we bring together diverse areas – different functions as well as external partners. Pooling from an unrivalled source of knowledge, we share, learn and challenge. It powers us to decode business needs and apply our technical know-how to add greater value. Rise to the challenge of shaping the future of an evolving business in the technology space.About The RoleThe Enterprise AI Platforms & Services Team are responsible for building and running the platforms, tooling and infrastructure that powers AstraZeneca’s ambition to use AI in every step of the value chain, from discovering new compounds to patient safety systems.We are looking for a Senior AI Operations Engineer to be a hands‑on technical leader within our AI/ML platform operations function. Reporting to the AI for Operations Lead, you will be the primary executor and technical driver across our growing platform estate, built on Azure and AWS. The adecuado candidate will have deep, current experience in site reliability engineering and will combine strong individual contribution with mentoring and technical leadership of L1 and L2 engineers.You will operate within a three-tier operations structure (L1 Runbook Operators, L2 Site Reliability Engineers, L3 Product Engineering interface) and be the senior hands‑on practitioner who builds the automation, designs the observability, and drives the continuous improvement flywheel day-to-day. You will be a key contributor to AI-augmented operations — building and refining AI-powered tools for runbook querying, incident pattern analysis,
and automated diagnosis.This is not a traditional support engineer role. This is a senior technical position for someone who sees operations as an engineering discipline and who can design, build, and instrument systems to the highest standard while coaching others to do the same.Key AccountabilitiesTechnical Execution and DesignDesign, build, and maintain the centralised observability layer — instrumentation, dashboards, alerting rules, and telemetry pipelines using platforms such as Datadog, New Relic, Grafana, or SplunkArchitect and implement AI-augmented operations tooling: conversational runbook interfaces, AI-driven incident analysis, and pattern recognition systemsLead complex incident response, perform root‑cause analysis, and produce actionable post‑mortemsContribute patches and instrumentation to product engineering codebases, ensuring they are architecturally sound and address root causesContribute to vendor evaluation to assess and recommend appropriate technologies and automation tooling.Automation and Continuous ImprovementBuild and maintain the automation that powers the continuous improvement cycle: every incident results in a runbook, an automation, or a patchOwn the technical execution of automation investments prioritised by frequency, resolution time, and blast radiusProactively identify and eliminate toil, maintaining the standard of no more than 50% of SRE time on reactive workDevelop and maintain runbooks, ensuring they are accurate, current, and progressively automatedOperational Readiness and Platform OnboardingImplement the Operational Readiness Gate — assess platforms against defined criteria before they transition from product engineering to operations ownershipShape platform operability during development — contributing instrumentation, health checks, and observability hooks before platforms enter BAUProvide technical input to platform architecture reviews from an operational perspectiveMentoring and Team ContributionMentor and coach L1 operators and junior L2 SREs, building their technical capability and engineering mindsetIdentify operators with engineering aptitude and support their development pathway from L1 to L2Contribute to the operations review by preparing flywheel metrics: L1 resolution rate, repeat incident rate, automation coverage, and toil budget complianceFoster a culture where SREs are engineers whose product is operational excellenceStakeholder CollaborationPartner with product engineering teams on post‑mortems, handover assessments, and operational obligation deliveryRepresent operational requirements in technical design discussionsCommunicate incident findings, operational risks, and improvement proposals clearly to the Operations Lead and wider teamEssentialCANDIDATE KNOWLEDGE, SKILLS AND EXPERIENCEBSc/MSc degree in Computer Science or related quantitative or analytical fieldSignificant hands‑on experience as a Site Reliability Engineer or platform operations engineer at scale — you build and run systems, not just direct othersStrong expertise in observability platforms (Datadog, New Relic,
Grafana, Splunk, or equivalent) including dashboard design, alerting strategies, and telemetry pipeline implementationStrong working knowledge of OpenTelemetry, distributed tracing, and structured logging standardsProven track record of designing and implementing automation that materially reduces operational toil — strong scripting skills in Python and Bash, with the ability to build robust tooling beyond one‑off scriptsExperience assessing platform readiness and contributing to operational handover processesStrong hands‑on skills with cloud infrastructure (Azure and/or AWS) including container orchestration, serverless architectures, and managed servicesExperience running or contributing significantly to post‑mortem processes and translating findings into preventive engineering workAbility to implement precise technical solutions — you can instrument a system, configure meaningful alerts, and build automation that intervenes at the right pointExperience mentoring junior engineers and contributing to team developmentDesirableExperience applying AI/ML to operational challenges — intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfacesFamiliarity with the AstraZeneca technology estate or regulated pharmaceutical environmentsExperience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)ITIL, SRE, or operational excellence certifications or equivalent practical frameworksExperience working within multi‑tier support structures with clear escalation pathsInfrastructure‑as‑code expertise (Terraform, CloudFormation)Awareness to GxP and audit trail awareness.Personal QualitiesAn engineering mindset applied to operations — you see every repeated manual task as an automation opportunityComfort operating at the boundary between deep technical work and collaborative team contributionCreative, collaborative, and resilientStrong communicator who can explain technical complexity clearly to both engineers and leadershipA bias toward systems thinking — you address root causes, not symptomsSelf‑directed with the ability to manage competing priorities across multiple platformsWhen we put unexpected teams in the same room, we unleash bold thinking with the power to inspire life-changing medicines. In‑person working gives us the platform we need to connect, work at pace and challenge perceptions. That’s why we work, on average, a minimum of three days per week from the office. But that doesn’t mean we’re not flexible. We balance the expectation of being in the office while respecting individual flexibility. Join us in our unique and ambitious world.#EAIDate Posted17‑ago‑2026Closing Date31‑ago‑2026AstraZeneca embraces diversity and equality of opportunity. We are committed to building an inclusive and diverse team representing all backgrounds, with as wide a range of perspectives as possible, and harnessing industry‑leading skills. We believe that the more inclusive we are, the better our work is. We welcome and consider applications to join our team from all qualified candidates, regardless of their characteristics. We comply with all applicable laws and regulations on non‑discrimination in employment (and recruitment), as well as work authorization and employment eligibility verification requirements.
#J-18808-Ljbffr
📌 Senior AI Operations Engineer (Barcelona)
🏢 AstraZeneca
📍 Barcelona