09 oct
|
Personio
|
Madrid
Your taskscplace is the platform for project and portfolio management that leading companies use to steer their most complex initiatives – grown in the DACH region and, with our launch in the US, on its way to becoming an international provider. As a Senior SRE in our Cloud Operations team, you will run our existing cplace Cloud 1.0 reliably and securely for customers in automotive, life sciences, manufacturing and retail – and actively shape the architecture and operating model of our Kubernetes-based cplace Cloud 2.0. AI is in cplace’s DNA: we use it intensively across all our work and expect you to apply it productively and critically and to help us take it further.- cplace Cloud 2.0: Co-building the Kubernetes platform – from cluster design, networking and storage to tenant isolation and scaling – plus planning and driving the migration of customer environments from Cloud 1.0- Everything as code: Development of reusable Terraform modules, GitOps repositories and our own platform software (self-service portal, APIs, automation), including code reviews, automated tests and policy as code- cplace Cloud 1.0: Operation and improvement of our environment of Linux servers, containers, SQL databases and Elasticsearch; lasting resolution of bugs, findings and capacity issues (incident and problem management, post-mortems); conversion of manual procedures into versioned, tested code (Ansible, Terraform, n8n)- Improvement of monitoring, logging and alerting; ownership of backup & recovery, disaster recovery and business continuity, including regular testing- Technical implementation of security and compliance requirements (e.G. SOC 2, ISO 27001, GDPR) and optimisation of cost and capacity- Driving AI in operations, e.G.
for incident and log analysis and agents for runbooks and routine tasks- Close collaboration with product development for smooth releases, technical representation of the team in customer meetings, tenders and customer projects, and knowledge sharing through internal and external documentation and mentoring- On-call duty in a fair rotationWhat we expect- Degree in computer science or a related STEM field, or comparable vocational training, plus several years (ideally 5+) of experience as an SRE, DevOps or platform engineer in business-critical production environments- Solid hands-on experience with Kubernetes in production (operations, upgrades, troubleshooting, storage, networking) and with at least one cloud provider – AWS, GCP, Azure and/or Hetzner Cloud a strong plus- Deep experience with infrastructure as code (Terraform, Ansible), CI/CD, GitOps and Git-based collaboration via pull requests and code reviews – with modular, tested code that stays maintainable for the team- Strong Linux skills and experience with SQL databases (e.G. MariaDB), Elasticsearch/OpenSearch and observability stacks (e.G. Prometheus, Grafana, Loki/ELK)- Solid software engineering skills, ideally in Go or Python – tools and automation with tests and clean structure rather than one-off scripts; confident Bash scripting a given- Good understanding of cloud security (e.G. network segmentation, secrets management, WAF)- Hands-on experience with AI tools in everyday engineering and a good sense of their strengths and limits- Customer-focused, structured way of working, ability to explain technical topics clearly, and fluent German (at least C1) and EnglishThis is what we offer- Adaptable work model and remote work- Room for creativity, co-designand further development- Competitive salary, 30days vacation& sabbaticaloption- Hardwareofyourchoice#J-18808-Ljbffr
📌 Senior Site Reliability Engineer (F/M/D) - Spain (Remote)
🏢 Personio
📍 Madrid