Skip to the content.

Are your Grafana and Prometheus alerts all green while the service quietly degrades? Most alerting stacks aren't broken. They're desensitized. Send me your stack and I'll come back with the handful of alerts that actually matter, plus one thing you almost certainly have wrong. Around two hours of my time, free. If it's useful we can talk about going deeper; if not, you keep the analysis.

Get your free alert audit

or email me directly if you would rather not fill in a form.

What I need from you

Five answers. They take about three minutes to write, and they’re what makes the audit worth anything:

  1. Name and company.
  2. Your stack. Grafana? Prometheus? Loki? Are SLOs defined? Self-hosted or Grafana Cloud?
  3. Roughly how many alerts does on-call receive per week? An estimate is fine. This is the single most useful number you can give me.
  4. Current state, in one sentence. “We’re firefighting.” / “We don’t know what to monitor.” / “Everything is green but things still break.”
  5. A link to a public dashboard. Optional, and it unlocks an extra 30 minutes of free review.

If your stack has no Grafana, no Prometheus and no SLOs, this audit won’t help you, and I’ll tell you that instead of wasting your afternoon.

What you get

Two to three hours of work, delivered as a short written analysis:

What you don’t get

This part is deliberate, and I’d rather be blunt about it up front:

What happens after

Usually the audit surfaces two problems: one I can describe in a paragraph, and one that needs real debugging to confirm its actual impact. The second is the work I charge for: a fixed-scope engagement that ends with the root cause fixed, explicit no_data handling, corrected rules, and an alerting plan your on-call will actually run.

No pressure either way. If the free analysis is all you wanted, it’s yours, and you owe me nothing.

Send me your stack

or email me directly if you would rather not fill in a form.