- Plan, manage, and oversee all aspects of a Production Environment
- Define strategies for Application Performance Monitoring and optimization in the Prod environment
- Respond to Incidents and improvise a platform based on feedback and measure the reduction of incidents over time.
- Support the deployment of code into multiple lower environments. Support current processes with an emphasis on automating everything as soon as possible.
- Design, develop, and standardize the monitoring and Alerting mechanism for the supported applications.
- Take a holistic approach to problem-solving by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize the mean time to recover.
- Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation, and refinement.
- Analyze ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
- Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
- Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
- Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.
- Work with a integral team spread across tech hubs in multiple geographies and time zones.
- Ability to share knowledge and explain processes and procedures to others.
- Share knowledge and mentor junior resources
- Able to perform on-call duties on a rotational basis.
- Occasional off-hours work required.
- Candidate should incline Training and should be a good trainer and ready to mentor others
Required Skills
- Linux
- Shell Scripting
- ITIL / ITSM
- SQL - Basic / Good to have
- Application Troubleshooting
- Ansible/Chef (Basic)
- Any Monitoring tool (Preferred Splunk/Dynatrace)
- Jenkins - CI/CD
- Groovy Scripting/YAML
- Git basic/bit bucket
Nice to Have
- Payments domain experience (Payment Flows, Switching, Authorization, Settlements)
- Event-Driven Architecture experience
- Exposure to distributed or cloud-native environments.
📌 Site reliability engineer (España)
🏢 Intuition It – Intuitive Technology Recruitment
📍 España
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.