Role summary
Improve reliability through clear service objectives, observability, incident learning and platform automation.
What you will do
- Define and review SLIs, SLOs and reliability risks
- Build monitoring and incident-response tooling
- Partner with service teams on resilience
- Reduce operational toil through automation
What you need
- Production operations experience
- Strong Linux, cloud and scripting skills
- Observability and incident-management knowledge
- Systems-thinking approach
Preferred experience
- Multi-region systems
- Chaos or resilience testing
- Platform engineering experience
What the employer offers
- Remote-first collaboration
- Home-office support
- Learning budget
- Structured on-call rotation
Interview process
- 01
Application review
- 02
Reliability interview
- 03
Systems case
- 04
Leadership conversation
- 05
Decision
