05 sep
|
Psd Group
|
Barcelona
05 sep
Psd Group
Barcelona
Senior Infrastructure & Storage Engineer
Barcelona (Hybrid)
Our client is a general leader in aviation technology, delivering innovative IT and communications solutions that connect airlines, airports, aircraft and governments worldwide. The Senior Infrastructure & Storage Engineer owns the implementation, integration, testing, and engineering lifecycle of on-premises compute, Linux, distributed storage, and related data-centre infrastructure.
This is a senior, hands-on role with particular emphasis on bare-metal infrastructure, Ceph, MinIO, Linux, Kubernetes storage integration, capacity, backup and recovery, and operational knowledge transfer. The engineer works closely with Platform Architecture, Operations, Kubernetes & Platform Engineering, Network Engineering, Security, vendors, and delivery teams
Build, configure, troubleshoot, and lifecycle-manage physical and virtual servers across heterogeneous data-centre environments.
• Engineer and support CentOS, RHEL, or equivalent Linux platforms, including operatingsystem, service, package, certificate, disk, memory, process, and performance issues.
• Execute data-centre refreshes, hardware deployments, firmware and operating-system changes, and replacement activities in accordance with approved designs.
• Validate compute, network, storage, power, capacity, resilience, monitoring, and support dependencies before production acceptance.
• Support selected Windows infrastructure where required.
Distributed Storage and Data Protection
• Administer and troubleshoot Ceph, including cluster health, OSDs, monitors, placement groups, pools, capacity, performance, replication, and failure domains.
• Engineer and support distributed MinIO deployments, object-storage availability, capacity, healing, replication, and recovery.
• Diagnose storage latency, degraded redundancy, disk and node failures, data-path issues, and capacity risks using an evidence-based approach.
• Plan, document, and test backup, restoration, disaster recovery, and data-protection procedures; Kubernetes Infrastructure Integration
• Support bare-metal Kubernetes nodes and their operating-system, hardware, network, containerruntime, and storage dependencies.
• Engineer and troubleshoot Kubernetes persistent storage, CSI drivers, StorageClasses, persistent volumes, mounts, and Ceph-backed workloads.
• Collaborate with the Kubernetes & Platform Engineer during cluster upgrades, node maintenance, capacity changes, recovery testing, and stateful workload incidents.
• Understand core Kubernetes concepts sufficiently to diagnose whether failures originate in the workload, node, network, CSI, storage, or underlying infrastructure.
Automate repeatable infrastructure provisioning and configuration using Terraform, Ansible, Bash, Python, or equivalent tools.
• Implement Prometheus metrics, dashboards, and actionable alerts for hosts, hardware, Ceph, MinIO, storage paths, capacity, and backup health.
• Integrate infrastructure logs and telemetry with platforms such as New Relic and Elasticsearch where appropriate.
Testing, Handover, and Engineering Support
• Plan and execute functional, integration, resilience, failover, capacity, backup, restoration, upgrade,
and rollback testing.
• Provide Level 3 support, lead root-cause analysis, and implement permanent corrective actions for complex infrastructure incidents.
• Advanced Linux systems administration and troubleshooting capability.
• Strong experience with storage systems and data-protection concepts, including replication, quorum, failure domains, capacity, performance, backup, and recovery.
• Experience working with physical servers, disks, controllers/HBAs, firmware, operating systems, virtualization, and data-centre dependencies.
• Experience automating infrastructure through Ansible, Terraform, Bash, Python, or equivalent technologies.
• Experience implementing monitoring, alerting, capacity management, and operational procedures.
• Strong incident troubleshooting, root-cause analysis, risk management, and vendor escalation skills.
• Direct administration of Ceph in production, including recovery from degraded states and performance investigations.
• Distributed MinIO on dedicated servers or bare-metal infrastructure.
• Kubernetes, Rancher, RKE2/RKE, CSI, and persistent-volume integration.
• Prometheus, New Relic, Elasticsearch, or comparable observability platforms.
• MariaDB, PostgreSQL, or other infrastructure database administration.
• CentOS/RHEL, virtualization, selected Windows infrastructure, and hybrid Azure integration.
• Business-continuity and disaster-recovery exercises across multiple data centres.
Candidates do not need to be application-platform or cloud specialists. Depth in Linux, on-premises infrastructure, distributed storage, safe recovery, and operational knowledge transfer is more important than matching every desirable product.
📌 Senior Infrastructure Engineer - IT Services (Barcelona)
🏢 Psd Group
📍 Barcelona