Act as incident commander leading response to incidents, owning proactive detection, and holding the rollback trigger during degraded releases. Run the daily proactive trend review, chair post-incident reviews.
Make an Impact By:
- Serve as incident commander for all incidents: run the bridge, enforce escalation clocks, own communications, and drive verify-and-close.
- Run the daily proactive trend review to surface degradation before it alerts; maintain the watch list.
- Execute and coordinate rollback decisions for degraded releases and invoke the contact-centre fallback where required.
- Tune the safeguard set and alert thresholds; drive alert precision and noise-budget discipline.
- Chair post-incident reviews and own corrective-action tracking and runbook updates.
- Mentor members of the team when needed; own shift-handover standards.
Skills for success:
- Bachelor’s or Master’s degree in Computer Science or a related field
- 6+ years in production operations
- Experience in customer-facing production environments
- Monitoring/observability and log/trace analysis
- Incident management, SLOs, and runbook operations
- Understanding of AI agent failure modes and routing
- Analytical and pragmatic, with the ability to interpret governance principles into implementation plans
- Clear communicator who can explain complex technical risks and solutions to non-technical stakeholders
- Self-driven and proactive, comfortable working in a fast-paced environment