There is a familiar failure mode in security tooling: a scanner or middleware that inspects traffic thoroughly and correctly, and adds two hundred milliseconds to every request while doing it. The security team is satisfied. The product team is not, because their p99 latency just got worse for every user, forever, in exchange for protection most requests never needed in the first place.
The instinct to build security analysis directly into the request path is understandable. If you want to inspect a request, the request is right there, mid-flight, already being handled. Why not look at it on the way through? The answer is that the moment you put analysis logic between a client and its response, you have made a promise you cannot always keep: that whatever you are checking will finish fast enough not to matter. Rule sets grow. Checks get more sophisticated. The promise gets harder to keep every quarter, and nobody notices until a demo day, or a traffic spike, turns two hundred milliseconds into two seconds.
Treat latency as a hard constraint, not a target
The useful reframe is to treat added latency as a hard budget set at close to zero, not a soft target you try to keep low. A target gets negotiated against every time a new rule seems worth the cost. A hard constraint forces the actual architectural question up front: if the response cannot wait on the analysis, where does the analysis happen instead?
The general answer is some version of the same pattern used across most non-blocking systems: let the response go, and do the expensive work on a copy, off the request's own timeline. The client gets its answer at normal speed. The analysis still happens, just not in a place where its cost is charged to someone waiting on a page to load. This is not a novel idea. It is the same reasoning behind async logging, behind webhook delivery queues, behind almost anything labelled "fire and forget" done properly. The discipline is in applying it consistently to security tooling specifically, which has a cultural habit of treating synchronous, in-path checks as more rigorous, when in practice they are just more visible when they go wrong.
Non-blocking does not mean unaccountable
The objection to this pattern is usually some version of: if the check happens after the response is already gone, what stops something bad from getting through in the gap? The honest answer is nothing stops it in that single request. What the pattern buys you is that the gap is measured in the time it takes the analysis to run, typically a small number of seconds, not the time it takes a human to notice a problem exists. A finding that surfaces moments after the fact and gets acted on quickly is a completely different risk profile from either blocking every request forever, or not checking at all.
This tradeoff is not appropriate for every kind of check. Authentication has to be synchronous; you cannot let a request through and revoke it after the fact. The distinction that matters is between checks that gate access and checks that detect and report a problem. Gating has to be fast and in-path by necessity. Detection does not, and pretending otherwise is how detection systems end up slow enough that someone eventually disables them under load, which is worse than never having built them.
This is the standard we hold ourselves to
We think about this constraint explicitly whenever we are asked to sit anything in front of a client's live traffic, AuditGate included. The rule we apply is simple to state and genuinely uncomfortable to enforce once a rule set starts growing: whatever we add to inspect traffic must never become the reason a response is slow. That discipline is a permanent design constraint, not a launch-day promise we plan to relax once the product is established. The moment the exception feels reasonable is exactly the moment the discipline stops meaning anything.
Written by
Akintola Stephen Iyanu, founder and engineer at Zynterra. More about the studio.