06 sep
|
Psd Group
|
Barcelona
06 sep
Psd Group
Barcelona
Lead Cloud Platform Engineer Barcelona (Hybrid) Our client is a general leader in aviation technology, delivering innovative IT and communications solutions that connect airlines, airports, aircraft and governments worldwide.
The Lead Cloud Platform
Engineer owns the implementation, integration, testing, and engineering lifecycle of cloud infrastructure, managed Kubernetes, and cloud-native platform services. The role translates approved architecture into secure, reliable, scalable, observable, and operationally supportable cloud platforms. This is a senior, hands-on engineering role.
It combines Azure platform engineering, Kubernetes, infrastructure as code, cloud networking and identity, observability, reliability, and structured transition to Operations. The engineer works closely with Platform Architecture, Operations, Infrastructure & Storage Engineering, Network Engineering, Security, vendors, and Product and Delivery teams.
Cloud Platform
Engineering • Build and operate Azure cloud infrastructure and shared platform services in accordance with approved architecture, security, resilience, and governance standards.
- Engineer cloud networking, identity, access controls, secrets, compute, storage, load balancing, private connectivity, and service integration.
- Implement reusable platform patterns consistently across development, test, preproduction, and production, including hybrid integration where required.
- Monitor cloud capacity, availability, performance, consumption, and cost; Kubernetes and Cloud-Native Engineering • Build, administer, upgrade, scale, and troubleshoot production Kubernetes platforms, with emphasis on Azure Kubernetes Service or comparable managed services.
- Diagnose issues involving worker nodes, container runtimes, CNI, CSI, CoreDNS, ingress, Services, scheduling, workload performance, and cloud dependencies.
- Implement resilient workload patterns using Deployments, StatefulSets, DaemonSets,
health probes, requests and limits, disruption budgets, affinity, and topology-spread constraints.
- Plan and execute cluster lifecycle changes with compatibility checks, staged validation, stop/go gates, and tested recovery procedures. Build reusable cloud infrastructure and platform modules using Terraform, Bicep, or equivalent infrastructure-as-code tools.
- Develop supporting automation using Bash, Python, PowerShell, or Go and reduce undocumented manual procedures.
- Implement policy, security, configuration, and deployment controls as code wherever practical. Observability and Performance • Implement cloud, Kubernetes, infrastructure, and service metrics, dashboards, and actionable alerts using Azure Monitor, Prometheus, or equivalent tooling.
- Use application performance monitoring and distributed tracing platforms such as New Relic to correlate platform and application behaviour.
- Integrate logging platforms such as Elasticsearch and establish clear ownership, thresholds, escalation paths, and runbooks.
- Troubleshoot application latency across DNS, cloud load balancers, gateways, ingress, Services, pods, network paths, managed services, databases, and storage dependencies. Testing and Production Readiness • Plan and execute functional, integration, capacity, resilience, failover, upgrade, rollback, backup, and restoration testing.
- Validate the complete system, including application, network, compute, storage, monitoring, security, and operational dependencies.
- Provide technical leadership, mentoring,
and reviews of platform implementations, automation, test plans, and operational documentation.
- Collaborate across Architecture, Operations, Infrastructure & Storage, Network Engineering, Security, FinOps, vendors, and delivery teams. Strong recent, hands-on experience engineering Azure or comparable public-cloud platforms.
- Strong production experience administering and troubleshooting Kubernetes, preferably a managed service such as AKS.
- Practical knowledge of cloud networking, identity and access management, security controls, compute, storage, load balancing, private connectivity, and DNS.
- Strong Linux fundamentals and practical knowledge of Kubernetes networking, ingress, scheduling, workload controllers, storage integration, and container runtimes.
- Strong experience with infrastructure as code and automation, preferably Terraform, Bicep, or equivalent technologies.
- Experience with Git, CI/CD, GitOps, and controlled production change practices.
- Experience using metrics, logs, tracing, and alerts to resolve cross-layer incidents, and implementing cloud resilience, recovery, governance, and cost controls.
- Strong troubleshooting, root-cause analysis, risk management, and technical documentation skills.
- Azure Kubernetes Service, Azure networking, Azure identity, Azure Policy, Key Vault, Azure Monitor, and related platform services.
- Prometheus, New Relic, Elasticsearch, and Argo CD or comparable observability and GitOps tooling.
- Cloud security posture management, policy as code, landing-zone implementation, and regulated environments.
- FinOps practices, capacity optimization, cloud cost controls, and hybrid connectivity.
- Business-continuity, disaster-recovery, and production resilience testing. Strong cloud, Kubernetes, automation, security, and operational fundamentals are more important than matching every product listed.
📌 Cloud and Platform Lead Engineer (Barcelona)
🏢 Psd Group
📍 Barcelona