01 ago
|
AstraZeneca
|
Barcelona
01 ago
AstraZeneca
Barcelona
Senior AI Operations Engineer – AstraZeneca Spain
Location: Barcelona, Spain (or Madrid). Visa sponsorship available. Full‑time contract.
Overview
AstraZeneca’s Enterprise AI Platforms & Services Team seeks a senior technical leader to build and run the platforms, tooling and infrastructure that power AI across the business. The role focuses on platform operations on Azure and AWS, with a strong emphasis on site reliability, observability, automation and AI‑augmented tools.
Responsibilities
- Design, build and maintain the centralized observability layer – instrumentation, dashboards, alerting rules and telemetry pipelines.
- Architect and implement AI‑augmented operations tooling such as conversational runbook interfaces, AI‑driven incident analysis and pattern recognition systems.
- Lead complex incident response, perform root‑cause analysis and produce actionable post‑mortems.
- Contribute patches and instrumentation to product engineering codebases.
- Evaluate vendors and recommend technologies and automation tooling.
- Build and maintain automation that powers the continuous improvement cycle; create runbooks, automation or patches for every incident.
- Own the technical execution of automation investments prioritised by frequency, resolution time and blast radius.
- Maintain the standard of no more than 50% of site‑reliability‑engineer time on reactive work.
- Develop and maintain runbooks, ensuring they are accurate, current and progressively automated.
- Implement the Operational Readiness Gate, assess platforms against defined criteria before transition to operations ownership.
- Shape platform operability during development – instrumentation, health checks and observability hooks before BAU.
- Provide technical input to platform architecture reviews from an operational perspective.
- Mentor and coach L1 operators and junior L2 SREs; support their development pathway from L1 to L2.
- Prepare flywheel metrics: L1 resolution rate, repeat incident rate, automation coverage, and toil budget compliance.
- Collaborate with product engineering teams on post‑mortems, handover assessments, and operational delivery.
- Represent operational requirements in design discussions; communicate incident findings, risks and improvement proposals to leadership.
Candidate Knowledge, Skills and Experience
Essential
- BSc/MSc in Computer Science or related quantitative field.
- Significant hands‑on experience as a Site Reliability Engineer or platform operations engineer at scale.
- Expertise in observability platforms (Datadog, New Relic, Grafana, Splunk or equivalent).
- Working knowledge of OpenTelemetry, distributed tracing and structured logging standards.
- Track record of designing and implementing automation that reduces operational toil; proficient in Python and Bash.
- Hands‑on skills with cloud infrastructure (Azure and/or AWS) including container orchestration, serverless architectures and managed services.
- Experience running or contributing to post‑mortem processes and translating findings into preventive engineering work.
- Ability to implement precise technical solutions – instrumentation, meaningful alerts, and automation that intervenes at the right point.
- Experience mentoring junior engineers and contributing to team development.
Desirable
- Experience applying AI/ML to operational challenges – intelligent alerting, automated diagnosis, predictive incident detection or conversational interfaces.
- Familiarity with AstraZeneca technology estate or regulated pharmaceutical environments.
- Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines).
- ITIL, SRE or operational excellence certifications or equivalent frameworks.
- Experience in multi‑tier support structures with clear escalation paths.
- Infrastructure‑as‑code expertise (Terraform, CloudFormation).
- Awareness of GxP and audit trail requirements.
Personal Qualities
- Engineering mindset applied to operations; sees repetitive manual tasks as automation opportunities.
- Comfort operating between deep technical work and collaborative team contribution.
- Creative, collaborative, resilient.
- Strong communicator – explains technical complexity clearly to engineers and leadership.
- Systems thinking – addresses root causes, not symptoms.
- Self‑directed and able to manage competing priorities across multiple platforms.
Work Model
In‑person attendance: minimum three days per week from the office with flexibility for remote work when appropriate.
Equal Opportunity Statement
AstraZeneca embraces diversity and equality of opportunity. We are committed to building an inclusive and diverse team representing all backgrounds, with a wide range of perspectives, and harnessing industry‑leading skills. We welcome and consider applications from all qualified candidates, regardless of their characteristics. We comply with all applicable laws and regulations on non‑discrimination in employment, as well as work authorization and employment eligibility verification requirements.
#J-18808-Ljbffr
📌 Senior AI Operations Engineer at AstraZeneca (Barcelona)
🏢 AstraZeneca
📍 Barcelona