Estimated reading time : 3 min · Published October 2, 2026
Choose the journey to protect
Describe the essential action, audience and dependencies that can interrupt it. A business site may need visitors to read an offer and send enquiries; a shop needs correctly recorded orders. Separate page loading, action completion and subsequent handling because these steps can fail differently.
The Google SRE Workbook describes service objectives based on indicators and agreement between stakeholders. Adapt the approach to the site’s size. A small team can start with one journey, an observable indicator and an owner rather than many dashboards without associated decisions.
Define events and observation windows
Write what counts as an attempt and success, where it is observed and which exclusions are justified. Deliberately invalid input differs from loss of a valid enquiry. Synthetic monitoring and actual usage observe different populations; keep their results separate.
Define the reading window and minimum useful evidence. Percentages are unstable with few events on a low-traffic site. Show attempt and failure counts alongside the rate and investigate reproducible cases. Avoid collecting unnecessary messages or personal information to measure a technical outcome.
Choose an objective before making a commitment
An internal objective describes a desired level and guides decisions. A contractual commitment has scope and consequences requiring separate agreement. Avoid adopting a high percentage merely because it sounds reassuring. Compare visitor needs, operating constraints and recovery costs.
In a fictional calculation, 10 failures among 1,000 attempts produce 99% success. This does not describe duration, affected people or severity: identical counts can represent minor inconvenience or lost orders. Time-based availability requires a different definition; do not convert indicators without explaining the method.
Connect thresholds to actions
Decide who investigates alerts, what evidence establishes an incident and how affected people are informed. A confirmed outage may require a workaround or rollback. An isolated alert may need verification before escalation. Reduce uncertainty and response time rather than generating messages for every fluctuation.
Document quality declines during changes. The error-budget policy chapter describes agreed responses to reliability gaps. A small site might simply delay an optional change until a critical defect is resolved without claiming to implement a large service’s operating model.
Verify observers and review incidents
Test monitoring itself: does it distinguish page availability from completed actions? Can it detect a lost notification? Does it issue an understandable alert that receives a response? Disruptive scenarios belong in authorised environments. Production checks must avoid fictional orders and accidental messages.
After incidents, connect timelines, impact, fixes and regression checks. Revise indicators if faults escaped detection. The incident-response guide and acceptance log help assign actions and retain findings. Useful objectives evolve with journeys and observed risks.
Reference documents
Content updated on October 2, 2026
