Increase the reliability and operational clarity of mission services through automation, observability, and disciplined incident response.
What you'll do
- Define SLOs and build actionable telemetry across services and platforms.
- Automate toil and strengthen capacity, resilience, and recovery practices.
- Facilitate blameless incident reviews and drive durable corrective action.
What helps
- 5+ years operating distributed production systems.
- Hands-on Kubernetes, cloud, monitoring, and scripting experience.
- Comfort participating in a sustainable on-call rotation.