What to demand from a platform.
A vendor-neutral checklist for the North Star
Draft, under review.
Use this when choosing an event-management platform, or when checking whether the one you have can take you towards the North Star. It names no vendors: every platform is measured against the same directives.
A platform is only half the answer. Sponsorship, one yardstick and planning for people are the business’s own work; for those directives the question is what the platform must support.
Requirements, directive by directive
Signal
The platform must
- Correlate alerts with each other and with documented runbook scenarios to decide whether an incident is actionable.
- Apply one response policy across the business, not a rule per manager.
Test it in a demo
Send a lone CPU alert and a CPU alert that coincides with failing checkouts; show which one becomes actionable, and why.
The platform must
- Group related signals and point to a likely root cause before an incident or problem ticket is raised.
- Measure time to root cause, and how it changes when the same issue returns.
Test it in a demo
Replay a past outage and show how quickly the platform points at its root cause, compared with the first time.
The platform must
- Take in approved change records and pipeline deploys as context.
- Mark alerts inside an approved change window, record them, and keep them out of incident management.
- Link each incident to the change that came before it, approved or not.
Test it in a demo
Deploy a change outside the change process and show where that deploy appears next to the alerts it caused.
The platform must
- Carry expected impact, response time, what the service delivers, customer or internal scope, and the availability, functionality or performance lens on every alert, under your own field names or mapped from each team’s.
- Build relationships from every data source, each marked with where it came from and how far it can be trusted.
- Feed what it learns back towards the CMDB.
Test it in a demo
Show one alert from a team with its own naming and one from a team on your standard, and how both end up with the same impact hints.
The platform must
- Track each monitoring gap found after an outage until the new monitoring is live.
- Replay the outage against the new monitoring to prove it would now be caught.
Test it in a demo
Take a gap from a recent outage and show the replay that proves the new monitor fires in time.
Response
The platform must
- Run an approved remediation automatically when an incident matches its root cause.
- Record in the trigger which approval authorised it.
- Give each candidate fix a confidence score.
Test it in a demo
Show an automated fix and, from the monitor itself, who approved it and when.
The platform must
- Import the fixes your vendors document into a catalogue with confidence scores.
- Let each part of the business adopt a fix as it is, or with its own extra steps.
Test it in a demo
Pick a known issue from one of your technologies and show its documented fix in the catalogue, scored and ready to approve.
The platform must
- Forward agreed priorities to one escalation service.
- Send the responder and the service owner the notice each needs, and tell people about an automatic fix only when the business was affected.
Test it in a demo
Show the notice a responder gets and the one a service owner gets for the same incident.
Proof
The platform must
- Capture outages so they can be replayed.
- Replay without paging real responders, while still opening an incident for the record.
Test it in a demo
Replay last month’s biggest outage and show what would have paged, and what would have been fixed automatically.
The platform must
- Watch its own pipeline: heartbeats from every source, delivery time, and sources that go quiet.
- Give service owners a report proving their monitoring works in normal running, during blackouts and in disaster-recovery tests.
Test it in a demo
Silence one alert source and show how quickly the platform notices.
The platform must
- Close incidents from the monitor’s own recovery signal.
- Verify fixes for log-based alerts against the log or the system of record.
Test it in a demo
Show a fix for a log-based alert and the evidence that the error stopped.
The platform must
- Run replays as repeatable tests for AI agents before they act.
Test it in a demo
Show an AI agent tested against a replayed outage, and its result.
The platform must
- Enforce the limits your security team sets on automation and AI, and record every action for audit.
Test it in a demo
Ask an agent to do something outside its limits and show the refusal in the audit record.
Organisation
The platform must support
- Report every line of business with the same metric definitions.
Test it in a demo
Show two lines of business side by side on the same eight metrics.
The platform must support
- Approve automation risk for many services at once, with reporting senior leaders can read.
Test it in a demo
Show one approval covering a whole group of services, and the report a sponsor would see.
The platform must support
- Show which teams already automate their fixes, and let their processes be reused.
Test it in a demo
Show the teams with automated fixes, and one team’s runbook reused by another.
The platform must support
- Report the time automation frees, per team, so the people can be planned before it happens.
Test it in a demo
Show how much on-call time automation saved one team last quarter.
The two-axis checklist
Score each directive twice: how well it is done today, and how much it matters to the business. Where both are far apart, you have found your priorities.
The six levels
- 0 Absent Technical maturity: Not in place. Business value: No effect on the business.
- 1 Initial Technical maturity: Done by hand, in a few places, when someone remembers. Business value: A benefit someone can describe but not measure.
- 2 Emerging Technical maturity: Some teams rely on it; others don’t. Business value: A benefit you can point to in one area.
- 3 Established Technical maturity: The normal way of working, and measured. Business value: A measured effect on SLAs, downtime or cost.
- 4 Advanced Technical maturity: Every line of business works this way, with the numbers to show it. Business value: Measured effects across the business, reported regularly.
- 5 Optimised Technical maturity: Improves itself: learns from every incident and replay. Business value: A proven return that leadership reviews and funds.
Scores of 3 and above need evidence: a dashboard, a config, a runbook, a replay result or a report.
| Directive | Technical maturity (0–5) | Business value (0–5) | Evidence |
| 1 Actionable means business impact | | | |
| 2 Root cause first | | | |
| 3 Know what changed | | | |
| 4 Impact travels in the alert | | | |
| 5 Close every monitoring gap | | | |
| 6 Approve once, automate after | | | |
| 7 Use what vendors already solved | | | |
| 8 One escalation path | | | |
| 9 Prove it by replay | | | |
| 10 Prove the watch is still standing | | | |
| 11 Trust is verified, not claimed | | | |
| 12 Replays are the evals for AI | | | |
| 13 Guardrails belong to security | | | |
| 14 One yardstick | | | |
| 15 Sponsorship sets the ceiling | | | |
| 16 Start with those already automating | | | |
| 17 Plan the people before the automation | | | |
The full Signal Assessment starts from a discovery template: the same directives, scored in your own words.
Know where you stand before the next storm.
Thirty minutes, free and specific. We look at your estate, and I tell you honestly whether I’m the right pilot.
Plot a course: book a 30-minute call