21 sep
|
Factorial Hr
|
Barcelona
21 sep
Factorial Hr
Barcelona
Overview
As a Senior Platform Engineer, you own the CI/CD Kubernetes platform end to end, shaping capacity, reliability, upgrades, security, and cost. You work with a cross-functional engineering team to keep a self-hosted runner fleet fast and scalable for hundreds of developers. You’ll drive performance improvements, manage bare-metal infrastructure, and turn incident learnings into robust runbooks and guardrails. This role blends hands-on infrastructure with product-minded collaboration to make builds faster and more predictable.
Compensaciones / Incentivos
- Alan private health insurance
- Wellhub fitness facilities
- Cobee expense savings
- Language classes
- Office breakfast and organic fruit
- Pet-friendly office
Responsabilidades
- Rightsize runner tiers using weekly CPU/memory data and update deployment via pull requests
- Investigate and fix flaky CI jobs by tracing issues to container runtime, kernel, or clock drift
- Add and onboard machines to the runner fleet with provisioning and networking
- Reduce queue wait times by identifying blockers across pipeline stages
- Upgrade clusters or controllers without disrupting users
- Collaborate with product engineers to optimize workflows that are slow for a reason
- Own incident response for the platform, run blameless postmortems, and implement alerts and runbooks
- Operate bare metal infrastructure with focus on Linux, networking, and troubleshooting
Requisitos principales
- 5+ years of production infrastructure with Kubernetes at the core
- Deep knowledge of Kubernetes scheduling, resources, evictions, node pressure, DaemonSets, controllers and operators
- Strong Linux and container fundamentals (containerd or Docker, cgroups, storage, networking)
- Terraform and infrastructure-as-code, with modular design and state hygiene
- CI/CD at platform level; GitHub Actions experience including self-hosted runners
- Scripting in Bash and at least one of Python or Go
- Observability with metrics, logs, traces; OpenTelemetry and Prometheus tooling
- Operational maturity: on-call, incident command, postmortems
- Capacity and cost awareness; ability to size a fleet and explain billing
- Clear written English and documentation skills
- Collaborative mindset
- Problem-solving and ownership
- Communication and ability to translate tech details for engineers
- Kubernetes architecture and operations (scheduling, limits, evictions, DaemonSets, operators)
- Bare-metal provisioning and troubleshooting
- GitHub Actions runner platform and runner controller concepts
📌 Senior DevOps Engineer, CI/CD Platform (Developer Experience) - barcelona
🏢 Factorial Hr
📍 Barcelona