Are your Grafana and Prometheus alerts all green while the service quietly degrades? Most alerting stacks aren't broken. They're desensitized. Send me your stack and I'll come back with the handful of alerts that actually matter, plus one thing you almost certainly have wrong. Around two hours of my time, free. If it's useful we can talk about going deeper; if not, you keep the analysis.
or email me directly if you would rather not fill in a form.
What I need from you
Five answers. They take about three minutes to write, and they’re what makes the audit worth anything:
- Name and company.
- Your stack. Grafana? Prometheus? Loki? Are SLOs defined? Self-hosted or Grafana Cloud?
- Roughly how many alerts does on-call receive per week? An estimate is fine. This is the single most useful number you can give me.
- Current state, in one sentence. “We’re firefighting.” / “We don’t know what to monitor.” / “Everything is green but things still break.”
- A link to a public dashboard. Optional, and it unlocks an extra 30 minutes of free review.
If your stack has no Grafana, no Prometheus and no SLOs, this audit won’t help you, and I’ll tell you that instead of wasting your afternoon.
What you get
Two to three hours of work, delivered as a short written analysis:
- A prioritized list of the 3–5 alerts that actually matter in your stack: the name, why it matters, and what ignoring it costs you.
- One sharp finding, verified. Typical shapes: an alert that stays green because
no_datawas never configured; a rule evaluating against a log store the logs never reached; a datasource that silently resolves to the wrong backend; a dashboard panel whose query no longer matches the label it claims to group by. - A real count: N alerts per week → N that deserve a human.
What you don’t get
This part is deliberate, and I’d rather be blunt about it up front:
- I don’t apply the fix. The finding is stated, not implemented.
- I don’t hand over the full noise-reduction plan.
- I don’t touch anything in production.
What happens after
Usually the audit surfaces two problems: one I can describe in a paragraph, and one that needs real debugging to confirm its actual impact. The second is the work I charge for: a fixed-scope engagement that ends with the root cause fixed, explicit no_data handling, corrected rules, and an alerting plan your on-call will actually run.
No pressure either way. If the free analysis is all you wanted, it’s yours, and you owe me nothing.
or email me directly if you would rather not fill in a form.