AIO APEX

Semantic Cache Policy Engine

A Semantic Cache Policy Engine is an ACCEPT/REJECT gate that decides whether a cached LLM answer is safe to reuse, instead of trusting embedding similarity alone. The term was introduced by Mohammad Safari in 2026.

What is a Semantic Cache Policy Engine?

Semantic caching has an obvious appeal: users ask the same questions in different words, and every re-answered question is GPU money spent twice. The standard approach embeds the incoming prompt, finds the nearest cached prompt by cosine similarity, and serves the cached answer above a threshold. It works until it serves a confidently wrong answer. In a compliance-sensitive product one wrong answer costs more than a year of cache savings.

How it differs from a gateway semantic cache policy

A Semantic Cache Policy Engine is not a gateway cache policy. Gateway policies such as Zuplo's Semantic Cache Policy, Apigee's SemanticCacheLookup, and Azure API Management's llm-semantic-cache-store decide when to cache by a similarity threshold. The engine decides whether a specific cached answer is safe to return, by entailment, and rejects contradictions that similarity cannot see. It can sit behind any of those gateways as the safety gate.

The negation trap

To an embedding model, “allow the user to access the account” and “block the user from accessing the account” are nearly identical: same domain, same entities, same structure, cosine similarity above any threshold you would set. To your customer they are opposites. Similarity measures topic; safety depends on meaning.

How the engine decides

Stage one is deliberately permissive: dense retrieval with a small embedding model and a fixed random projection shortlists cached candidates cheaply, tuned for recall. Stage two makes the decision: a natural-language-inference cross-encoder reads the two prompts in both directions and applies a strict policy. Accept only if neither direction contradicts and at least one direction entails. Contradiction in either direction is an instant REJECT. ACCEPT and REJECT are the only outputs.

Both models are public and off the shelf. Stage one is BAAI/bge-small-en-v1.5, whose 384-dimension output is projected to 128 dimensions through a fixed seeded Johnson–Lindenstrauss matrix, gated at cosine 0.70. Stage two is cross-encoder/nli-deberta-v3-small, run in both directions. Together they are about 270 MB.

Stage one shortlists cached candidates by dense retrieval. Stage two reads both prompts in both directions with a natural-language-inference cross-encoder and returns ACCEPT or REJECT.IncomingpromptStage onedense retrieval,tuned for recallStage twobidirectional NLIcross-encoderACCEPTREJECT
The two-stage gate. Stage one shortlists cheaply; stage two decides.

Measured results

On 959 cross-verified adversarial pairs, including 220 pairs engineered to flip meaning while keeping the words nearly identical, the engine returned 0 false approvals on all 220 negation pairs. Across the whole set it made 9 false approvals in 959 pairs, and every one of them was a hard conditional inversion — 9 of the 133 pairs in that category, about 6.8% of it, and none at all in the other three.

ConfigurationPrecisionRecallF1TP / FP / FN / TN
Shipped (pre-filter on)0.9750.9910.983349 / 9 / 3 / 598
Two-stage only (pre-filter off)0.9720.9940.983350 / 10 / 2 / 597

Shipped configuration: precision 0.975 (95% CI 0.953–0.987, Wilson), recall 0.991 (95% CI 0.975–0.997). One ACCEPT/REJECT decision per candidate pair. Nothing is trained, so there is no train/test split: both models are off-the-shelf and frozen.

False approvals by category, shipped configuration — the residual error is concentrated in one place rather than spread across the set:

Negation0 / 220
Entity / number swap0 / 154
Unrelated0 / 100
Conditional inversion9 / 133

Stage one alone is not a safety gate, which is the point of the second stage. Dense retrieval on its own scores precision 0.412, recall 1.000, F1 0.584 on the same pairs: it accepts nearly everything it retrieves. Adding the inference stage and the pre-filter cuts false approvals from 502 to 9.

Measured on one NVIDIA A100 80GB PCIe against a live vLLM backend over 539 replayed requests across five datasets, the median decision latency per dataset ranges 16.9–28.5 ms and averages 23.4 ms. That average is the mean of the five per-dataset medians, not a pooled median over all 539 requests. The p95 reaches 206 ms on one dataset. Declining a reuse adds 1.5–5.6 ms to a cache miss. The same path on CPU is 100–200 ms, not 23 ms.

Dataset: 1,008 candidate pairs were generated by claude-sonnet-5 and then independently label-verified by a model from a different provider, OpenAI gpt-5.6-sol; 959 pairs survived cross-provider agreement, a 95% keep rate. Every misclassification was then manually inspected and confirmed to be a genuine model limit rather than label noise. All figures come from that corpus and from synthetic benchmarks. There is no production or customer validation.

When not to use it

It is not built for personalised or state-dependent queries such as “what's my balance?”, where the correct answer changes between identical questions. Those bypass the cache by policy. A deterministic pre-filter runs ahead of the semantic stages, rejecting cross-tenant reuse, requester- or time-specific queries, and numeric disagreement, for example “$75,000” versus “$750,000”. That pre-filter is pattern matching, not a solver, so the bypass is a policy rather than a guarantee. Hard conditional inversions remain the documented residual error: 9 of 959 pairs, and all nine of the system's false approvals.

Who it is for

Teams running LLM inference in the UK where a wrong cached answer is a compliance event. Contact via LinkedIn: https://www.linkedin.com/in/mosafariuk

Reference implementation: https://github.com/mosafariuk/semantic-cache-policy-engine