Benchmarks
What the policy layer costs, how we measured it, and how to measure it yourself on your own account.
Policy-layer overhead
Gateco sits between an AI system and the data it retrieves, so the honest question is what the permission check adds on top of the vector query that would have happened anyway. That is the number below. It is the service-side policy layer only, not end-to-end retrieval latency, and it excludes query embedding.
Those are the 2026-08-16 run. We publish it as the headline figure because it is the most conservative of the five: across every recorded run the p50 ranges from 10.0ms to 16.0ms and the p95 from 13.55ms to 26.2ms. The runs at the high end carry stall outliers many times their own p99, which is host contention on a shared development machine rather than a property of the software. We publish those runs too.
Conditions
Every figure on this page comes from the following setup. It is a development machine, not a tuned production deployment, and the backend ran with verbose logging enabled, which makes the overhead numbers conservative rather than flattering.
- Connector
- pgvector, local
- Corpus
- 1,000 vectors, 1,536 dimensions, 100 registered resources
- Policy
- 1 active RBAC policy with 5 rules
- Metadata resolution
- sidecar
- Requests
- 150 measured, 10 warmup, paced at 8 rps, top_k 10
- Backend
- Development backend with DEBUG enabled
Artifacts
Every run behind the figures above, directly fetchable.
- policy-overhead-bench-2026-08-16.json
The headline run: full percentile tables for all four series, plus the methodology block describing the exact conditions.
- policy-overhead-bench-all-runs.json
All five recorded runs, including the two that came out worst. Identical code and methodology across every run.
Measure it yourself
The numbers above are ours, measured on our machine. You do not have to take them on trust. Every Gateco account, including the free one, has a Performance Self-Test page that measures the same thing against the running service and reports it back to you.
It reads the policy-layer overhead already recorded on your own retrievals, which is the same quantity published on this page, and it runs the policy engine live against your own active policies. Neither measurement calls a vector database or an embedding provider, so running it does not consume your retrieval quota.
Create a free account and run it, or sign in and open Performance Self-Test.
Why this page exists
An earlier version of this site carried an overhead claim we had not actually measured. We benchmarked it, found real bugs while doing so, corrected the copy, and added a CI check that blocks the unverifiable phrasing from coming back. The write-up is at Auditing our own latency claim.