What do you do when the failure already happened and you can’t make it happen again? You don’t reproduce it. You reconstruct it from what was recorded, which is the reason to record things before you need them.
Three records, joined by the trace ID from Chapter 7, tell the story of one request:
trace_id 7c1e9a52-...
request POST /api/licenses consumer 4 504 10.2 s
outgoing POST statamic/sites timeout
exception ConnectionException warning
The request log and the exception are already there. The middle line is the one teams forget: a record of every call you make to a provider. Laravel’s HTTP client fires an event for each response, so one listener covers every driver you will ever write:
// app/Providers/AppServiceProvider.php, in boot()
Event::listen(function (ResponseReceived $event): void {
$request = $event->request;
Log::info('provider.response', [
'method' => $request->method(),
'host' => parse_url($request->url(), PHP_URL_HOST),
'status' => $event->response->status(),
]);
});
Because the trace ID is in Context, it is attached to this log line without the listener knowing it exists. A ConnectionFailed event covers the calls that never got an answer. If you run Nightwatch, it already records outgoing calls and you can skip the listener.
Log the host and the status, not the URL and the body. A provider URL can contain an identifier you shouldn’t keep, and a provider’s response body is somebody else’s data.
Jobs get the same treatment for free. Context travels with a dispatched job, so the log lines written by DeleteLicense ten minutes later carry the trace ID of the request that queued it. When a consumer quotes an ID from the X-Trace-Id header, one search returns the request, the job, every provider call either of them made, and the exception.
Objectives
Four signals tell you what is happening. They don’t tell you whether it is acceptable. Is a 95th percentile of 800 ms good? Is one failed request in five hundred?
Without an answer agreed in advance, every graph is an argument. One person sees a blip, another sees an incident, and the decision goes to whoever is most worried that day.
A service level objective is that answer, written down. It has three parts: a measurement, a target, and a window.
- Availability: 99.5% of requests, over 30 days, answer without a 5xx that is this API’s fault.
- Latency: 95% of
GET /api/licensesrequests answer within 500 ms. - Background work: 99% of license deletions complete within five minutes of the request.
Three decisions hide in those sentences, and each is worth a moment’s thought.
What counts as a failure. A 422 is the consumer’s mistake and doesn’t count. A 504 because the provider was down is harder. It isn’t your code’s fault, and your consumer doesn’t care whose fault it is. I count it, because an objective that excuses your dependencies measures your comfort and not your consumers’ experience.
What the target is. Not 100%. A target of 100% can’t be met, so it can’t guide anything. 99.5% over thirty days allows about three and a half hours of failure, and that allowance is the useful part.
What you do with the allowance. The gap between the target and perfection is an error budget. While there is budget left, ship: deploy on Friday, try the risky migration. When it is spent, stop shipping features and spend the time on reliability until it recovers. That turns “should we slow down?” from a debate into a reading.
And it changes what you alert on. Don’t page when one request fails. Page when the budget is burning fast enough to run out, which is the definition of a symptom that matters.
You don’t need new tooling for this. The request log already holds every status and duration. An objective is a query over it, run on a schedule, with a number that goes on the dashboard beside the four signals.
The First Five Minutes of an Incident
When something is down, the first five minutes decide whether the next hour is calm or chaotic. Do these in order, before you open a single file.
- Check the health route on every instance, not only the one that alerted. It tells you at once whether the problem is this application’s own dependencies.
- Look at the four signals for the last thirty minutes, not the last five. An error rate that started creeping twenty-five minutes ago is your real starting point, not the moment the alert fired.
- Read the errors by status code. A wall of 504s means the provider. A wall of 500s means you. A wall of 429s means one consumer, or your own limiter set too low.
- If it is scoped to one consumer, pull a trace. The outgoing-call line usually shows which call failed before you have read any application code.
- Only then open a file. Everything before this step is reading data you already collected. Everything after it is debugging, and debugging without the first four steps is guessing with extra confidence.
Chapter 13 Summary
Health checks:
- Extend
/upwith aDiagnosingHealthlistener that checks what a request needs: the database and the cache. - Never check a provider there. A dependency’s outage must not take your servers out of rotation.
Signals:
- Volume, latency, errors, and saturation, visible without reading a stack trace.
- Latency at the 95th percentile. Errors by status code, never added together.
Alerting:
- Page on symptoms a consumer can see, not on retries still within budget.
- Sort exceptions by severity before they reach the error tracker.
Tracing:
- Record every provider call, by host and status, with one listener.
- With a trace ID in
Context, the request, its jobs, its provider calls, and its exceptions read as one story.
Objectives:
- Write down what counts as a failure, a target below 100%, and a window. An objective is a query over the request log.
- Spend the error budget on shipping, and stop shipping when it is gone.
During an incident:
- Health route, four signals over thirty minutes, errors by status code, one trace. Only then open a file.