22 sep
|
Factorial Hr
|
Madrid
22 sep
Factorial Hr
Madrid
Hey there, Kubernetes people! Hundreds of developers push to a large Ruby on Rails monorepo, and every push lands on a runner fleet that the Developer Experience team builds, runs and keeps fast. That fleet is self-hosted Kubernetes on bare metal in European datacentres, in the low hundreds of machines, running episodic GitHub Actions runners that peak above a thousand concurrent pods.
We own all of it: the hardware, the clusters, the runner images, the caches and the delivery pipelines on top. We add nodes by hand rather than letting an autoscaler do it, so capacity planning is genuinely part of the job. Shave a minute off the average build and every engineer in the company gets that minute back, several times a day.
Own the CI and CD Kubernetes clusters end to end: capacity, reliability, upgrades, security and cost. Run the self-hosted GitHub Actions platform at scale with Actions Runner Controller. Runner scale sets, episodic pods, docker-in-docker, and the runner images themselves. Keep the data services our tests depend on fast and healthy. MySQL, Redis and ClickHouse come up per job, alongside Rails application containers.
Design and tune the caching that makes builds quick: node-local overlay, image and artifact caches we build ourselves, plus pull-through registry mirrors. Bring queue wait and build duration down, with measurement behind it. Manage it all as code.
Terraform planned and applied from pull requests, GitOps delivery with Flux and Argo CD, Kustomize and Helm for manifests. Provision and operate bare metal. Linux, networking across datacentre segments, and troubleshooting that sometimes ends up at the disk or the kernel.
Rightsizing runner tiers against a week of CPU and memory data, then opening the pull request that changes them. Tracking a flaky job from a red check down to the container runtime, the kernel, or a clock that drifted.
Adding machines to the fleet: install, enrol, network, verify, document. Cutting p95 queue wait by working out which stage actually blocks.
5+ years running production infrastructure, with Kubernetes at the centre of it. ~ Real depth in Kubernetes: scheduling, requests and limits, evictions, node pressure, DaemonSets, controllers and operators.
Strong
Linux and container fundamentals. containerd or Docker internals, cgroups, namespaces, storage drivers, networking. ~ Terraform and infrastructure as code, with a feel for module design and state hygiene. ~ Comfortable with GitHub Actions, including self-hosted runners. ~ Bash, plus at least one of Python or Go. ~ Metrics, logs, traces, dashboards and alerts that people trust, with OpenTelemetry and Prometheus style tooling. ~ Clear written English.
Large
Ruby on Rails test suites, and what it takes to parallelise them honestly.
Operating
MySQL, Redis or ClickHouse yourself.
Bare metal: provisioning, hardware failure, and the different rhythm it has compared to cloud. Overlay and WireGuard style mesh networking, VLANs, and Kubernetes CNIs such as Cilium. Working across more than one cloud.
We use AWS, Azure and Cloudflare alongside our own hardware. Contributions to open source infrastructure projects. That's why our Engineering Team follows an office-first, adaptable approach. However, we also support remote work when it makes sense (20%) for deep focus or personal needs.
Intro Call: Chat with our Talent Partner about your journey and goals.
Final Coffee Chat: A relaxed conversation with our CTO & VP of Engineering to explore our vision, culture, and your growth. And that's it! The whole process is remote, using videoconferencing tools!
At Factorial, we're building the leading AI Business Management Software for companies of all sizes. Our platform centralizes key workflows across HR, Finance, Talent, Operations, and IT, freeing teams from manual processes so they can focus on what really matters: leading, growing, and taking care of their people.
We own it: We take responsibility for every project. We're dedicated to learning something new every day and, above all, share it. Alan as private health insurance Save expenses with Cobee Breakfast in the office and organic fruit
📌 Senior DevOps Engineer, CI/CD Platform (Developer Experience) (Madrid)
🏢 Factorial Hr
📍 Madrid