← All jobsALR-107

Alerian · Cloud & Platform

Site Reliability Engineer

Remote — U.S. · remotePublic Trust eligible

This opportunity is contingent on program demand or contract award.

The opportunity

Increase the reliability and operational clarity of mission services through automation, observability, and disciplined incident response.

What you'll do

  • Define SLOs and build actionable telemetry across services and platforms.
  • Automate toil and strengthen capacity, resilience, and recovery practices.
  • Facilitate blameless incident reviews and drive durable corrective action.

What helps

  • 5+ years operating distributed production systems.
  • Hands-on Kubernetes, cloud, monitoring, and scripting experience.
  • Comfort participating in a sustainable on-call rotation.