Skip to content
Cloud & platform engineering

Platform reliability & SRE enablement

Service level objectives, error budgets, on-call and runbooks, installed and then handed over.

6–8 weeksTypical duration

The problem this solves

You have dashboards and alerts, but nobody agrees on what 'up' means, and the on-call rotation is a list of people who happen to know things.

What you receive

Artefacts you can hold, and that you can accept or refuse — never a list of activities.

  • Service level objectives per critical service, agreed with the teams that own them
  • Alerting rewritten to fire on symptoms rather than on causes
  • An on-call practice with escalation, handover and a blameless review format
  • Runbooks for the incidents you actually have, written from your own history