If you’re passionate about technology, innovative projects, and making a real impact, your place is here.
We are a integral technology provider with 65+ years of experience, delivering a comprehensive portfolio of solutions across ID & Digital Government, Banking & Payments, and Trusted Connectivity. Within our Trusted Connectivity business unit, we develop cutting-edge solutions for the telecommunications industry—ranging from SIM cards and eSIMs to Subscription Management and secure connectivity services —connecting people, businesses, and devices worldwide.
We are looking for a highly analytical and business-oriented Site Reliability Engineer to design and enhance the reliability engineering architecture of our platforms, ensuring high availability, scalability, reliability, and observability through close collaboration with R&D;, DevOps, and Operations teams.
As a Site Reliability Engineer, you will be responsible for working mainly with Site Reliability Architect (SRA) and R&D; team to design resilient systems and operational processes that ensure the high availability, scalability, reliability and observability of our platforms.
Site Reliability Engineering
Ensure new features have been validate in terms of performance, reliability and saclability
Work together with DevOps team to improve existing and implement new, effective CI/CD processes.
Work together with Enablement engineer to produce automation tools needed for performance and reliability monitoring
Continuously evaluate and optimize system performance and capacity in order to maintain stable production platforms.
Identify, assess, and implement measures to eliminate potential risks that could impact the performance of systems and services.
Research, evaluate, test and advise at selecting appropriate new technologies or tools for improving site reliability
Observability & Monitoring
Monitor system performance, identifying bottlenecks, and execute pipeline optimization
Implement comprehensive service metrics to track and report on system reliability, performance, and efficiency.
Capacity Planning & Performance Engineering
Support in forecasting, scaling, and performance tuning.
Bachelor’s degree in Computer Engineering, Electronics Engineering, Telecommunications Engineering, or a related field.
~3+ years of experience in Site Reliability Engineering, Infrastructure Operations, DevOps, or a similar role.
~3+ years of experience within the telecommunications industry or related technology sectors.
~ Strong Linux administration skills (Red Hat, Ubuntu, or similar distributions).
~ Hands-on experience with cloud platforms, preferably AWS.
~ Experience designing and maintaining CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions, or similar).
~ Experience with monitoring and observability tools (Prometheus, Grafana, ELK Stack, Datadog, or equivalent).
~ Proficiency in scripting and automation using Python, Bash, or similar languages.
~ Experience with Infrastructure as Code (Terraform, Ansible, or equivalent).
~ Strong knowledge of containerization and orchestration technologies (Docker, Kubernetes).
~ Experience in performance monitoring, troubleshooting, and system optimization.
~ Experience working with SQL databases.
~ Advanced English communication skills (B2+/C1).
~ AWS, Linux, or Kubernetes certifications.
Knowledge of capacity planning and performance engineering.
Experience in telecom platforms, mobile services, or cloud-native architectures.
Join Valid and work on innovative, global technology projects within multicultural and multidisciplinary teams.
Flexibility: flexible working hours and remote work options to support work-life balance.
Well-being first: private medical insurance and life insurance.
We are committed to equal opportunities, free from discrimination concerning sex, age, race, sexual orientation, religion, education, social status, culture, or special needs such as illness or disability.
📌 SRE Engineer (Madrid)
🏢 Valid
📍 Madrid