19 sep
|
Ludium Lab
|
España
☁️ Are you passionate about Cloud infrastructure, automation and reliability and ready to take on a key role in a growing technology environment?
Join Ludium Lab, a technology company revolutionizing the cloud gaming industry! If you enjoy building reliable and scalable infrastructure, solving complex technical challenges and making systems more efficient and resilient, keep reading!
? ? Who are we?
Ludium Lab is a technology company founded in Barcelona, Spain , in 2012. We are specialists and a leading company in Cloud services and solutions , operating in more than 60 countries around the world. Currently, our work focuses on adapting our technology to areas such as cloud gaming platforms (developing Sora Stream), the automotive industry , metaverse solutions and SaaS/XR (VR/AR) .
With more than a decade of experience, our team develops high-quality and cost-efficient solutions based on virtualization and cloud streaming technologies .
? What will you do as a Senior SRE?
As a Senior Site Reliability Engineer , you will play a key role in ensuring the reliability, scalability, performance and availability of our Cloud infrastructure and services. You will work within the Cloud team, in close collaboration with the Development team , to automate processes, improve observability and make our systems more resilient.
? Your main responsibilities will be:
• Define and monitor SLIs and SLOs , establishing clear and measurable reliability objectives for our services.
• Manage Error Budgets and use reliability metrics to balance stability, development speed and product evolution.
• Automate infrastructure and operational processes using Infrastructure as Code .
• Work with containers and Kubernetes to operate and scale workloads in production.
• Implement and maintain monitoring, logging,
alerting and observability solutions, defining relevant metrics to measure the health of our services.
• Identify and resolve performance, availability and reliability issues across our systems.
• Participate in incident management , troubleshooting and root cause analysis.
• Lead and contribute to postmortems , identifying preventive actions and opportunities for improvement after incidents.
• Improve system resilience through automation, redundancy, capacity planning and proactive reliability practices .
• Identify opportunities to optimize Cloud resources, infrastructure costs and operational efficiency .
• Help establish and promote SRE best practices within the engineering organization.
? What are we looking for?
• Several years of experience working as an SRE, DevOps Engineer, Cloud Engineer or in a similar role , with a strong focus on reliability and systems operations.
• Experience defining and monitoring SLIs, SLOs and Error Budgets , using metrics to make decisions about service reliability.
• Solid experience working with at least one of the major Cloud platforms , preferably AWS.
• Strong knowledge of Linux systems, networking and infrastructure .
• Experience with Infrastructure as Code , preferably Terraform.
• Good knowledge of monitoring, logging, alerting and observability , using tools such as Prometheus, Grafana, ELK/OpenSearch or similar .
• Good scripting or programming skills, using Python,
Bash, Go or similar .
• Experience solving complex problems in production environments and participating in incident management.
- Ability to work autonomously, take ownership and collaborate effectively with multidisciplinary technical teams.
✨ It would be a plus to have:
• Experience with highly scalable, distributed or highly available systems .
• Experience working with cloud streaming, real-time services or platforms with demanding latency and availability requirements .
• Experience with GitOps or modern deployment practices .
• Experience in capacity planning and Cloud cost optimization .
• Previous experience in startups or fast-growing technology environments .
• Practical experience with Kubernetes and Docker in production environments will be valued.
? What do we offer?
• 100% remote work , with the possibility of using our offices in Barcelona and Madrid .
• Growth opportunities : at Ludium Lab, we are committed to your continuous development, both personally and professionally.
• Versátil working hours so you can manage your time efficiently.
• Flexible compensation plan and private health insurance .
• A collaborative environment, surrounded by a talented and multidisciplinary team.
• Ongoing support for your training and professional development .
• Openness to innovative ideas that challenge the status quo.
At Ludium Lab , we promote a culture of diversity, innovation and professional growth .
If you are passionate about Cloud infrastructure and reliability and want to contribute to building the technology that powers our next generation of Cloud services, we want to hear from you!
? Are you ready to join our team?
Your success will be our success.
📌 Senior Site Reliability Engineer (SRE) (España)
🏢 Ludium Lab
📍 España