17 ago
|
Palo Alto Networks
|
Madrid
17 ago
Palo Alto Networks
Madrid
Experteer Overview
In this role, you will join Team Engineering to operate and improve large-scale, multi-cloud production environments, ensuring reliable, secure services for tens of thousands of enterprise customers. You’ll work closely with cross-functional teams to reduce incidents and drive operational excellence in a high-availability, distributed setup. The role offers hands-on engineering in Kubernetes, Terraform, and Python, with a strong emphasis on automation, monitoring, and rapid incident response. This is a remote-friendly, fast-paced position at a market-leading cybersecurity company with a culture that values collaboration and impact.
Compensaciones / Beneficios
• Own and operate large-scale, integral production environments across multiple cloud providers (GCP, AWS, Azure)
• Monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
• Drive end-to-end troubleshooting across complex, distributed systems with high context switching
• Design, deploy, and improve monitoring and observability systems (Prometheus, Grafana)
• Collaborate with internal teams (CX, CS, Engineering) to ensure reliability and performance
• Work hands-on with DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
• Develop and maintain automation and tooling (primarily in Python)
• Gain deep understanding of system architecture and interconnected services
• Contribute to a culture of operational excellence in a high-scale, high-availability environment
• Champion asynchronous communication, documentation, and tooling standards for distributed teams
• On call responsibilities: Daytime hours 12:00-20:00 CET/CEST; occasional weekends and holidays by rotation
Responsabilidades
• 5+ years of experience in SRE roles in production environments at scale
• Strong hands-on experience with Kubernetes and Terraform
• Strong hands-on experience with at least one major cloud platform (GCP or AWS)
• Experience building and configuring monitoring systems (Prometheus, Grafana)
• Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
• Proficiency in Python for scripting and automation
• Proven success in a fully remote or distributed team environment
• Strong troubleshooting and problem-solving skills with incident handling
• Ability to work in fast-paced environments with high context switching
• Ownership-driven, proactive, and collaborative with strong communication skills
• Curious mindset and eagerness to learn
Requisitos principales
•
📌 Sr Staff Site Reliability Engineer (Madrid)
🏢 Palo Alto Networks
📍 Madrid