28 ago
|
Enfint
|
Barcelona
Описание: Codeway is a integral consumer tech company that develops and scales mobile apps across creativity, productivity, wellness, language learning, and entertainment. Its products serve more than 400 million users worldwide.
Задачи
- Define, instrument, and report on SLIs, SLOs, and error budgets across critical services
- Own observability end-to-end, including metrics, logs, traces, dashboards, and alerting
- Reduce alert noise and false positives
- Run reliability reviews and an error-budget policy
- Operate, scale, and upgrade the multi-cluster Kubernetes environment
- Act as the escalation point for cluster and platform issues
- Own capacity planning, performance, and cloud cost efficiency
- Build self-service platform tooling for product teams
- Embed security through RBAC, least privilege, secrets management, scanning, network policies, and patching
- Partner with the security function on vulnerability remediation, audit readiness, and secure‑by‑default infrastructure
- Own disaster recovery and validate RTO/RPO targets through drills and failure testing
- Contribute to architecture and production‑readiness reviews
- Lead the on‑call rotation and act as incident commander during production incidents
- Run blameless postmortems and track corrective actions through closure
- Build and maintain Infrastructure as Code with Terraform and CI/CD pipelines
- Enforce GitOps and progressive delivery with automated rollbacks
- Identify, measure, and eliminate operational toil through automation
- Establish and mature reliability capabilities during the first 12 months
- Establish shared security measurement and a common view of operational health across engineering and security
- Connect platform performance, reliability, delivery speed, security, and infrastructure cost to business operations
- Provide metrics that support company investment decisions
Требования
- 5–8 Years of experience in SRE, Platform,
or DevOps roles
- Experience operating high‑traffic, always‑on production systems at meaningful scale
- Hands‑on production Kubernetes experience, including upgrades, autoscaling, and troubleshooting under load
- Strong cloud engineering background
- Solid Linux and networking fundamentals
- Experience defining and operating with SLOs and error budgets
- Experience with Infrastructure as Code and CI/CD pipeline design
- Depth in observability tooling, instrumentation, dashboarding, and alert design
- Security‑first mindset with experience in least privilege, secrets hygiene, and vulnerability management
- Fluency in scripting and automation in at least one language
- Incident‑command experience, including on‑call ownership and blameless postmortems
- Ability to communicate clearly with engineers and leadership under pressure
- Nice to have: high‑scale consumer or mobile app backends, AI/ML inference workloads, GitOps, progressive delivery, service mesh, API gateways, multi‑region or multi‑cluster topologies, cloud cost optimization, FinOps, SOC 2, ISO 27001, GDPR, DevSecOps, chaos engineering, resilience testing, Kubernetes/cloud/DevOps certifications, supporting many independent services and teams
Условия
- Competitive compensation package
- Meal compensation
- Unlimited private health insurance and HPV vaccine coverage
- Pet adoption support covering primary healthcare expenses, parasite vaccinations, and microchip costs during the first year
- MacBook, iPhone 15 Pro, Magic Mouse, Magic Keyboard, adjustable desk, 4K screen, and other required gadgets
- Gym membership support
- Flexible schedule
- English course support
- Office located in central Barcelona at Edifici Estel
- Coffee shop, healthy snacks, free breakfast, and free lunch
- No dress code
- Gaming area with PS5
- Software subscription support
- Additional monthly compensation for public transportation
- Case study may be included in the recruiting process
#J-18808-Ljbffr
📌 sre engineer for high-availability services (Barcelona)
🏢 Enfint
📍 Barcelona