23 ago
|
Mirantis
|
Barcelona
23 ago
Mirantis
Barcelona
- We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-thinking cloud development and operations team
- In this role, you will contribute to the design, development, and operation of sophisticated cloud-based AI solutions built on the CNCF ecosystem including Kubernetes, running on cutting-edge hardware from leading vendors
- This role focuses mainly on deploying AI infrastructure built on NVIDIA-certified hardware, following architecture and implementation designs produced by our engineering team
- You will play a pivotal role in ensuring the reliability, security, and performance of container infrastructure, while mentoring team members and Mirantis customers to deliver high-quality software and services
- As a senior engineer, you will work closely with stakeholders to define technical strategies, solve complex challenges, and ensure the seamless integration of cloud and software services
- This is an excellent opportunity to make a significant impact while driving innovation in a rapidly evolving cloud ecosystem
- Work with geographically distributed international teams on technical challenges and process improvements
- Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software
- Collaborate with stakeholders to gather and refine technical requirements
- Optimize system performance, reliability, and scalability
- Troubleshoot, debug, and resolve complex technical issues
- Participate in code reviews to maintain high quality standards
- Stay up to date with industry trends and best practices in cloud operations and development
- Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance
- Facilitate knowledge transfer to customers during the delivery phases- A commitment to innovation, continuous learning, and delivering high-quality results
- Ability to travel up to 25% if needed, including internationally
- Excellent customer-facing communication skills
- Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams
- Excellent written and spoken English
- Experience with high-performance data center processing, networking, and storage
- At least 5 years of DevOps or Software Development experience or in a similar role
- Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines
- Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security
- Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight
- 5+ years of professional experience in DevOps, with a strong focus on cloud and infrastructure technologies, including Kubernetes and/or OpenStack
- Exposure to Golang and working knowledge of other programming languages (Python, JavaScript)
- Bachelor’s degree in Computer Science or a related field, or equivalent experience
- Extensive experience in network and/or storage architecture
- Experience with high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle or NVIDIA AI Enterprise
- Presence in the open source community including upstream contribution and conference presentations
- Prior experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware
📌 Senior Site Reliability Engineer (Barcelona)
🏢 Mirantis
📍 Barcelona