The North Star.
Read it all Space pauses · → next · ← back
Actionable incidents, to a person or an agent.
Leg 1 of 4 · Signal: What reaches anyone at all
1 Actionable means business impact
An incident is actionable when it is, or will be, hurting the business.
Noted, thank you
Leg 1 of 4 · Signal: What reaches anyone at all
2 Root cause first
Reduce the noise and find the root cause before incident or problem management starts.
Noted, thank you
Leg 1 of 4 · Signal: What reaches anyone at all
3 Know what changed
Every event carries what changed around it, approved or not.
Noted, thank you
Leg 1 of 4 · Signal: What reaches anyone at all
4 Impact travels in the alert
Every alert carries its own hints about business impact.
Noted, thank you
Leg 1 of 4 · Signal: What reaches anyone at all
5 Close every monitoring gap
A gap found after an outage stays open until the new monitoring is deployed and proven.
Noted, thank you
Leg 2 of 4 · Response: What happens next, and who acts
6 Approve once, automate after
Once the business approves how to fix a root cause, every later issue with that root cause is fixed automatically.
Noted, thank you
Leg 2 of 4 · Response: What happens next, and who acts
7 Use what vendors already solved
Catalogue every fix your vendors document, and score it against your own risk appetite.
Noted, thank you
Leg 2 of 4 · Response: What happens next, and who acts
8 One escalation path
Incidents reach people through one escalation service, and only when they should.
Noted, thank you
Leg 3 of 4 · Proof: How you know it works
9 Prove it by replay
Replay the outage to prove the fix works before it is needed.
Noted, thank you
Leg 3 of 4 · Proof: How you know it works
10 Prove the watch is still standing
Every service owner can prove their monitoring still works.
Noted, thank you
Leg 3 of 4 · Proof: How you know it works
11 Trust is verified, not claimed
An automated fix counts only when something independent shows it worked.
Noted, thank you
Leg 3 of 4 · Proof: How you know it works
12 Replays are the evals for AI
Every replay is also a test your AI agents must pass.
Noted, thank you
Leg 3 of 4 · Proof: How you know it works
13 Guardrails belong to security
Your security team sets what automation and AI may never do.
Noted, thank you
Leg 4 of 4 · Organisation: What makes it last
14 One yardstick
Every line of business measures event management the same way.
Noted, thank you
Leg 4 of 4 · Organisation: What makes it last
15 Sponsorship sets the ceiling
Event management is only as good as the most senior person who owns it.
Noted, thank you
Leg 4 of 4 · Organisation: What makes it last
16 Start with those already automating
Find who automates today, then give everyone else what they built.
Noted, thank you
Leg 4 of 4 · Organisation: What makes it last
17 Plan the people before the automation
Decide where the freed time goes before the first fix is automated.
Noted, thank you
How far are you from the star?
Actionable incidents, to a person or an agent.
Leg 1 of 4
Signal
Leg 2 of 4
Response
Leg 3 of 4
Proof
How far are you from the star? Take the self-assessment or book a call
Read it all
This is the ideal to steer by, not a claim: what matters is knowing how far you are from it and closing the distance step by step.
1Actionable means business impactAn incident is actionable when it is, or will be, hurting the business.
A CPU alert on its own is not actionable. It becomes actionable when other alerts correlate with it and show the impact, or when a documented runbook scenario already defines it as actionable and names the actions to run. Response rules set differently by every manager (everything investigated here, only P2 and above there) keep event management reactive for good.
Measured by Actionable ratio
Noted, thank you
2Root cause firstReduce the noise and find the root cause before incident or problem management starts.
Mean-time metrics are useful, but they are not the goal. The goal is to find the root cause faster every time, because a root cause with an approved remediation can be fixed automatically the next time it appears.
Measured by Time to root cause · Root-cause coverage
Noted, thank you
3Know what changedEvery event carries what changed around it, approved or not.
Most outages start with a change, so change context is the fastest way to the root cause, and it makes whoever writes the monitors, people or AI, think about change while they build them. Alerts inside an approved change window are expected: they are recorded in event management but not raised to incident management, and that record can show when an approved change hit a service its impact assessment did not name. Deploys made outside the change process are brought into view, and a business that runs ITIL reins them in. The hard part is the incentive: teams have to help connect their pipelines, and they rarely see the extra oversight as being on their side.
Measured by Time to root cause
Noted, thank you
4Impact travels in the alertEvery alert carries its own hints about business impact.
No enterprise keeps a complete, current service map, and some services are too short-lived to model at all. So the standard lives in the alert: the expected impact if the monitor fires, how the business should respond and within what time, what the service delivers, and whether customers or only internal users (payroll, HR) are affected. Each monitor also says whether it watches availability, functionality or performance. Owners who won’t update their CIs can be coached to write this into their alerts, and automation feeds it back into the CMDB over time. The standard has to work both ways, even inside one company: some stakeholders take it as principles and map it to their own naming, others want a fixed convention and change their data to fit, and the service supports both side by side. Relationships are the largest gap in every company, so every data source is examined for reliable ones, each marked with where it came from and how far it can be trusted.
Measured by Impact-hint coverage
Noted, thank you
5Close every monitoring gapA gap found after an outage stays open until the new monitoring is deployed and proven.
Most outages reveal something the monitoring missed. Tracking that gap until new monitoring is live, and replaying the outage to show it would now be caught, turns a business that reacts with hope into one that reacts once and is ready after that.
Measured by Monitoring gaps
Noted, thank you
6Approve once, automate afterOnce the business approves how to fix a root cause, every later issue with that root cause is fixed automatically.
The approval’s ID is built into the monitor’s trigger, so the monitor itself shows how its remediation was approved. Where several fixes exist (a stale runbook, clearer closure notes on a recent incident, different fixes for different parts of the business), each carries a direct and an indirect confidence score.
Measured by Automation coverage · Repeat incidents
Noted, thank you
7Use what vendors already solvedCatalogue every fix your vendors document, and score it against your own risk appetite.
Start from the contracts: every technology you pay support for comes with documented fixes. Collect them, catalogue them with confidence scores, and let each part of the business adopt them as they are or with its own extra steps. AI does the cataloguing; people stay in the discussion and give the final approval, but never become the bottleneck before that.
Measured by Automation coverage
Noted, thank you
8One escalation pathIncidents reach people through one escalation service, and only when they should.
Incident management forwards the agreed priorities to a single escalation service. Several overlapping notification tools cost money and create churn. The responder gets the actionable detail; the service owner gets the impact on their service and what automation did. A successful automatic fix informs people only if the business was affected. And a pilot that routes around the standard flow should be stopped, not declared a success.
Noted, thank you
9Prove it by replayReplay the outage to prove the fix works before it is needed.
Update the monitoring or the CMDB, trigger the issue again or reintroduce it synthetically, and watch the agreed steps run: a runbook fixes it, an incident is still opened so the record shows when it happened and how automation handled it, and the right people are informed. No real responder is paged unless the business wants to measure their response too.
Measured by Automation coverage · Monitoring gaps
- Outage
- Gap found
- Monitoring or fix updated
- Replay
- Proven
Noted, thank you
10Prove the watch is still standingEvery service owner can prove their monitoring still works.
Monitoring decays without anyone noticing: authors leave or move on, and the teams that inherit their monitors may not have the time or skills to understand them. So auditing monitoring is a routine activity that any product owner can rely on to answer one question: are our availability, functionality and performance goals still being met? A well-built, well-monitored service can show this in normal operation, during blackout and change periods, and in disaster-recovery events its owners take part in. The event pipeline’s own faults belong to SRE and first-line engineers, and once their cause is found, their fix is automated by default.
Measured by Monitoring gaps
Noted, thank you
11Trust is verified, not claimedAn automated fix counts only when something independent shows it worked.
For most monitoring, the tool’s own clear alert is the proof: it flows back through event management and closes the incident. Where there is no clear signal, as with many log-based checks, the fix must show that the error no longer appears in the log, or that another system of record is healthy again: the failed payroll workflow succeeded on retry and the systems downstream have what they need.
Noted, thank you
12Replays are the evals for AIEvery replay is also a test your AI agents must pass.
As AI takes over watching services, it needs the business’s own definition of a successful fix. Replays of real outages are exactly that: tests written in the business’s terms, run early and often.
Noted, thank you
13Guardrails belong to securityYour security team sets what automation and AI may never do.
AI agents have already been seen acting outside the limits they were given. Where a security team has no guidance of its own yet, published government guidance is the baseline.
Noted, thank you
14One yardstickEvery line of business measures event management the same way.
The same metrics and the same expectations, in a form that fits globally recognised processes. When one team owns the numbers and others quietly build their own, nobody is accountable to anyone.
Noted, thank you
15Sponsorship sets the ceilingEvent management is only as good as the most senior person who owns it.
It rarely reaches senior executives with enough urgency on its own. In a smaller business, that sponsor is easy to find. In a large one, win support across the organisation and help current management make the case urgent, until someone high enough owns the outcome. Approve automation risk at that level, and scale across a thousand services at once instead of one team at a time.
Measured by Staff per SLA met
Noted, thank you
16Start with those already automatingFind who automates today, then give everyone else what they built.
Some teams already automate their fixes. Learn how they won their sponsorship, partner with them on the larger effort, and hand their turnkey processes to the teams that don’t automate yet. Businesses agree on automation in principle and disagree on how much human involvement to keep, so the guide works with the processes already in place.
Noted, thank you
17Plan the people before the automation · Open problemDecide where the freed time goes before the first fix is automated.
A team budgeted for three people that automation could run with one has a manager who must now justify two roles. That friction quietly kills adoption. The best answer I know: each automation project starts with a plan for where its people can move, and the savings fund proactive retraining toward work that earns. I don’t claim this is solved.
Measured by Staff per SLA met
Noted, thank you
One yardstick
Eight headline metrics, the same in every line of business.
Example estate. Mean time to acknowledge, detect and resolve are still tracked, as supporting measures.
Where each directive sits in ITIL
For engineers trained on ITIL: the event-management activities, and the directives that serve each one.
Drafted from the published ITIL practice; under review.
| ITIL activity | Directives |
|---|---|
| Detection and notification of events | Know what changedImpact travels in the alertClose every monitoring gapProve the watch is still standing |
| Logging | Impact travels in the alert |
| Filtering | Actionable means business impact |
| Significance and classification | Actionable means business impactImpact travels in the alert |
| Correlation | Actionable means business impactRoot cause firstKnow what changed |
| Response selection | Approve once, automate afterUse what vendors already solvedOne escalation pathTrust is verified, not claimedReplays are the evals for AIGuardrails belong to security |
| Review and closure | Prove it by replayProve the watch is still standingOne yardstickSponsorship sets the ceilingStart with those already automatingPlan the people before the automation |
Know where you stand before the next storm.
Thirty minutes, free and specific. We look at your estate, and I tell you honestly whether I’m the right pilot.
Plot a course: book a 30-minute call