Zephiel API
Company15 October 20247 min read

Nine years of uptime reports and what they taught us about honesty

We have published every bad month since 2015. The commercial damage we feared has not once materialised.

Since 2017 our status page has shown per-API daily history that no human can edit, and since 2015 we have published monthly uptime per API including the bad ones. That decision was contested internally more than once. Nine years of data later, here is what it actually cost.

The fear

The argument against was straightforward and not stupid: a prospect comparing us to a competitor showing 100% will pick the competitor. We would be punished for measuring honestly while others were rewarded for measuring loosely.

What happened

In nine years we can identify two deals where published downtime was raised as an objection. Both closed. In both, the conversation moved from the number to what we did about it, and having an incident write-up to point at was worth more than a clean record would have been.

Meanwhile the number of deals where a prospect specifically cited the honest reporting as a reason for choosing us is much larger. We stopped counting properly around 2021, but it was in the dozens by then.

The asymmetry makes sense on reflection. Anyone technical evaluating a platform knows that 100% uptime over a year is not a real number. Publishing it does not build confidence in your reliability; it builds doubt about your measurement.

The internal effect we did not predict

The bigger benefit was inward. When downtime is published automatically and cannot be edited, the incentive to characterise an incident favourably disappears, because the characterisation does not change the bar on the chart.

That changed how our incident reviews went. We stopped spending the first twenty minutes negotiating whether something counted as an outage and started at what happened and what changes. Removing the ability to argue about the number removed the argument.

What we would tighten

Our definition of "degraded" was too generous for the first few years. An API returning errors for five per cent of requests was shown as degraded rather than down, and for the customer whose traffic was in that five per cent, it was down.

We tightened the thresholds in 2019 and our historical figures got worse as a result. We did not restate the earlier months, and the seam is visible in the data. That is the honest way to handle it.

Keep reading