Monitoring

Uptime, errors, traces and the pager at 3am.

0 products

Nothing matches that yet

Try a broader category, or add the product you were looking for.

Add a product

What this category covers

Knowing that something is broken before a customer explains it to you: uptime checks, error tracking, metrics, logs, traces and the alerting that wakes somebody up.

Products range from a single check that pings a URL, to platforms ingesting everything an infrastructure emits. The gap between them is mostly data volume and price.

Analysing customer behaviour belongs in analytics. Where the application runs is hosting and cloud.

Start from the outside

The most useful first purchase is the simplest: a check that loads your site from several places and tells you when it stops working.

External checks catch the failures internal dashboards miss, including expired certificates, broken records and a firewall change nobody announced. They are cheap, and for many small teams they cover most of what matters.

Error tracking comes second. Knowing which exceptions reach real users, how often, and since which release turns a vague report into a fix. Grouping quality is what separates products here, since a tool producing four thousand near-identical alerts is producing noise.

The four kinds of data

  • Uptime checks, from outside, on the paths customers actually use.
  • Errors, grouped, with the release and the user context attached.
  • Metrics, numeric and cheap to keep, for trends and thresholds.
  • Logs, detailed and expensive, for reconstructing what happened.
  • Traces, showing a request across services, for finding where time goes.

Most teams need the first two immediately, metrics soon after, and the rest when a system is complex enough that guessing has stopped working.

Retention is the quiet constraint. Comparing today with the same day last month needs a month of data, and entry plans frequently keep a week.

Alerting is a design problem

An alert that nobody acts on trains everybody to ignore the next one, and that habit is how outages get longer.

Alert on what users experience rather than on every internal threshold. A page loading slowly matters. A processor at eighty percent, on its own, usually does not.

Every alert needs three things: an owner, a documented action, and a route that reaches the owner when they are asleep. Anything failing that test should be a dashboard entry, not a notification.

Then review regularly. Alerts accumulate, conditions change, and a quarterly cull is the cheapest reliability work available.

Cost, and why bills grow without warning

Pricing runs by host, by user, by ingested volume, or by events, and the volume model is where surprises live.

A single service logging verbosely can multiply the bill while your traffic stays flat. Sampling, filtering at the source and shorter retention for detailed data are the usual controls, so ask what the product offers before you need them.

Custom metrics deserve a specific question, since they are frequently charged per unique series and a badly chosen label can create thousands of them. Ask for a spend cap rather than an alert, and look at what pricing pages hide for the rest of the pattern.

Running it yourself

Open implementations are capable and free to license. The operational cost is real: storage, retention, upgrades and somebody on call for the monitoring system itself.

That last point is the one to weigh. A self-hosted stack tends to fail in the same incident as the thing it was watching, which is exactly when you need it most.

Starting small without painting yourself in

A first monitoring setup should be boring and finished in an afternoon.

Two external checks on the paths that matter, error tracking wired into the application, and one alert route that reaches a person. That covers the majority of incidents a small product will have in its first year.

Add metrics when you start asking why something was slow rather than whether it was broken. Add traces when the system has enough moving parts that the answer is no longer obvious from a log.

The one decision worth making early is where data is collected and in what format. Instrumentation written against an open standard can be pointed at a different backend later, while instrumentation written against one vendor’s library becomes an argument against ever changing.

Keep a written record of what each alert means and who responds. That page is worth more at three in the morning than any dashboard.

Questions people ask

What should a small team monitor first?
Whether the site responds from outside your network, and whether errors are reaching users. Those two catch most incidents. Traces, profiling and dashboards can wait until there is something to investigate.
How is monitoring priced?
By hosts, by ingested data volume, by users, or by events. Data-volume pricing is the one that grows unexpectedly, because a single noisy service can multiply your bill without changing your traffic.
What retention do we need?
Long enough to compare against the same period last month, which usually means thirty days for detailed data and longer for aggregates. Entry plans often keep a week, which makes recurring problems invisible.
How do we stop alert fatigue?
Alert on symptoms users feel rather than on every threshold, give every alert an owner and a documented response, and delete alerts nobody acts on. An ignored alert is worse than no alert.
Is open-source monitoring cheaper?
The licence is free and the operation is not. Running your own storage, retention and on-call for the monitoring stack is real work, and it fails at the same time as everything else it watches.

Categories