04 ago
|
Tamarind Intelligence
|
Barcelona
04 ago
Tamarind Intelligence
Barcelona
ph3Overview /h3pWELCOME TO SITA /ppAt SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the integral air travel industry. /ppYou'll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting‑edge tech to keep operations running like clockwork. We don't just move the world forward-we're proud to be recognized as a bGreat Place to Work /b ® by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow. /ppAre you ready to love your job? /ppThe adventure begins right here, with you, at SITA. /ph3About The Role And The Team /h3pAs Lead Site Reliability Engineer/ Expert you will be responsible for the proactive support of products so that there is high product performance that is continuously improved. Responsible for identifying and resolving the root causes of operational incidents, implementing solutions to improve stability and prevent recurrence. Manages the creation and maintenance of the event catalog to trigger events and develop both manual remediation approaches and automated workflows to resolve alerts. Oversee the deployment of IT services and solutions ensuring successful integration with minimal disruption. Focuses on operational automation and integration to enhance efficiency and collaboration between development and operations within service operations. /ph3What You Will Do /h3ulliDefine, build, and maintain support systems to ensure high availability and performance. /liliHandle complex cases for the PSO. /liliImplement automation for system provisioning, self‑healing, auto‑recovery, deployment, and monitoring. /liliPerform incident response and root cause analysis (RCA) for critical system failures. /liliMonitor system performance and establish Service‑Level Indicators (SLIs) and Service‑Level Objectives (SLOs). /liliCollaborate with Development and Operations to integrate reliability best practices, including zero‑downtime architecture. /liliProactively identify and remediate performance issues. /liliWork closely with Product TE, ICE, and Service Architects for new product productization as SGS technical expert. /liliCoordinate with internal and external stakeholders to improve service performance and ensure high availability. /liliEnsure Operations readiness to support new products.
/liliAccountable within SGS for in‑scope product availability and performance. /li /ulh3Problem Management /h3ulliConduct thorough problem investigations and root cause analyses to diagnose recurring incidents and service disruptions. /liliCoordinate with Incident Management teams and collaborate with PSOs and Engineering/Product teams to implement permanent solutions. /liliMonitor effectiveness of problem resolution activities and provide regular reporting to ensure continuous improvement. /li /ulh3Event Management /h3ulliDefine, build, and maintain an event catalog specifying active events, thresholds, and remediation actions; optimize it for efficiency. /liliDevelop event response protocols, provide training, and ensure efficient incident handling. /li /ulh3Customer Operational Support /h3ulliCollaborate with Customer Success Managers to implement initiatives that enhance customer satisfaction and retention. /liliPrepare reports, documentation, and communication materials covering customer metrics, updates, and product changes. /liliIdentify and implement improvements in internal processes and workflows. /liliContribute to knowledge management resources such as FAQs and training materials. /li /ulh3Data Steward Responsibilities /h3ulliImplement data governance policies defined by the Data Owner and ensure adherence to standards. /liliMonitor data quality, consistency, and compliance on an ongoing basis. /liliAct as a Subject Matter Expert (SME) for data within the assigned area, providing guidance and answering queries. /li /ulh3Qualifications /h3h3ABOUT YOUR SKILLS /h3ulliBachelor's degree in Computer Science, Information Technology, Engineering, or a related field. /lili6+ years of experience in IT operations, service management, or infrastructure management, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Manager. /liliProven experience managing high‑availability systems and ensuring operational reliability. /liliExtensive experience in root cause analysis (RCA), incident management, and developing permanent solutions for recurring service disruptions. /liliHands‑on experience with CI/CD pipelines, automation, system performance monitoring, and infrastructure as code (IaC).
/liliStrong background in collaborating with cross‑functional teams (Development, Operations, Engineering, etc.) to improve operational processes and service delivery. /liliExperience managing deployments, conducting risk assessments, and optimizing event and problem management processes. /liliFamiliarity with cloud technologies, containerization, and scalable architectures, including zero‑downtime deployment strategies. /li /ulh3Technical Skills (Must‑to‑Have) /h3ulliStrong AKS On prem K8s skills and experience, /liliScripting (Ansible Bash, Python - combination of anything would be great), /liliAutomation, /liliCI/CD pipeline, /liliTerraform exposure, /liliAzure (or) AWS skill. /liliBasic DB skills. /liliStrong problem‑solving skills quick learner. /liliSRE mindset. /li /ulpbPlease note: /b Person will be working with global team, so he /she has to be flexible for overlap or stretch for operational urgency. /ph3What We Offer /h3ullibFlex Week: /b Work from home up to 2 days/week (depending on your team's needs) /lilibFlex Day: /b Make your workday suit your life and plans. /lilibFlex-Location: /b Take up to 30 days a year to work from any location in the world. /lilibEmployee Wellbeing: /b We have got you covered with our Employee Assistance Program (EAP), for you and your dependents 24/7, 365 days/year. We also offer Champion Health - a personalized platform that supports a range of wellbeing needs. /lilibProfessional Development: /b At SITA, we believe growth fuels innovation. Our learning ecosystem offers access to world‑class platforms and programs designed to help you thrive. From LinkedIn Learning, Microsoft's Enterprise Skills Initiative, and Airport Council International -available to all employees-to specialized solutions like Pluralsight for technology upskilling, Harvard Business Publishing for people leadership, Stanford for strategic development and many others, we align learning opportunities with your development plan and our business priorities. Your development journey is supported every step of the way. /lilibCompetitive Benefits: /b Competitive benefits that make sense with both your local market and employment status. /li /ulpSITA is an Equal Opportunity Employer. We value a diverse workforce. In support of our Employment Equity Program, we encourage women, aboriginal people, members of visible minorities, and/or persons with disabilities to apply and self-identify in the application process. /ph3Salary / Compensation Note /h3pHidden (-999) /p /p #J-18808-Ljbffr
📌 Lead Site Reliability Engineer/ Expert (Barcelona)
🏢 Tamarind Intelligence
📍 Barcelona