06 oct
|
HireHi
|
Cataluña
Описание:
Factorial provides AI for businesses that connects information about company operations and turns that context into action, supporting tasks such as candidate screening, expense management, answering questions, and report creation. It serves more than 16,000 companies across 112 countries.
Задачи:
Own the CI and CD Kubernetes clusters end to end, including capacity, reliability, upgrades, security, and cost; Run the self-hosted GitHub Actions platform at scale using Actions Runner Controller, runner scale sets, ephemeral pods, Docker-in-Docker, and runner images; Keep test data services fast and healthy, including MySQL, Redis, and ClickHouse running per job alongside Rails application containers; Design and tune build caching, including node-local overlay, custom image and artifact caches, and pull-through registry mirrors; Measure and reduce queue wait and build duration by instrumenting the platform, setting targets, and demonstrating improvements; Manage infrastructure as code using Terraform pull requests, GitOps delivery with Flux and Argo CD, and Kustomize and Helm manifests; Provision and operate bare metal, Linux, and networking across datacentre segments, troubleshooting down to disk or kernel level; Lead platform incident response, run blameless postmortems, and turn findings into alerts, guardrails, or runbooks; Treat the platform as a product for engineers by talking to users, observing where they get stuck, and building easy-to-follow paths; Right-size runner tiers using CPU and memory data and submit pull requests to change them; Trace flaky jobs from failed checks to the container runtime, kernel, or clock drift; Add fleet machines by installing, enrolling, networking, verifying, and documenting them; Reduce p95 queue wait by identifying the blocking stage; Upgrade clusters or controllers without disrupting users; Pair with product engineers to investigate slow workflows.
Требования:
- 5+ Years running production infrastructure, with Kubernetes at its centre
- Deep Kubernetes knowledge, including scheduling, requests and limits, evictions, node pressure, DaemonSets, controllers, and operators; experience debugging misleading cluster behavior
- Strong Linux and container fundamentals, including containerd or Docker internals, cgroups, namespaces, storage drivers, and networking
- Terraform and infrastructure-as-code experience, including module design and state hygiene
- Platform-level CI/CD experience with shared workflows, reusable templates, runner architecture, and build caching
- Experience with GitHub Actions, including self-hosted runners
- Bash scripting and experience with at least one of Python or Go
- Observability experience using metrics, logs, traces, dashboards, and alerts, with OpenTelemetry and Prometheus-style tooling
- Operational maturity, including on-call, incident command, postmortems, and root-cause investigation
- Capacity and cost awareness, including fleet sizing and explaining costs
- Clear written English and the ability to document systems
- Nice to have: large Ruby on Rails test suites and test parallelization, operating MySQL, Redis, or ClickHouse, bare-metal provisioning and hardware failure experience, overlay and WireGuard-style mesh networking, VLANs, Kubernetes CNIs such as Cilium, internal developer tooling, portals or platform APIs, supply-chain and build security, experience across multiple clouds, and open-source infrastructure contributions.
Условия:
- The office-first approach includes on-site work several days a week (80%) and remote work when appropriate (20%)
- Private health insurance with Alan
- Wellhub membership for gyms, pools, and outdoor classes
- Expense savings with Cobee
- Language classes
- Breakfast at the office and organic fruit
- Food discounts with Nora
- Pet-friendly environment
- Additional perks are shared during the interview process.
#J-18808-Ljbffr
📌 Devops engineer for ci/cd platforms (Cataluña)
🏢 HireHi
📍 Cataluña