Benchmark methodology

Evidence backed benchmarks.

A de-risked approach, backed by multiple reproducible benchmarks — each run head-to-head against standard RAG on the same model, the same questions, and the same retrieval.

Claim by claim

The metric breakdown.

Two measurement families back our data: an accuracy-and-efficiency study at 1,000-document scale on an 8B model, and a fleet-throughput ramp on a 30B model. We keep their attributions separate.

ClaimHow it was measured
91× less text per questionIn a 1,000-document test on an 8B model, Smart CAG re-read only the ~15-token question — where a standard RAG setup re-read ~1,352 tokens of document text, both given the same retrieved documents.
Same answer qualityAcross 500 questions, an independent — and more capable — model graded the answers for meaning and scored the two even: a statistical tie with RAG (details below).
faster first tokenIn the 30B fleet test with one request at a time, the first token came back in versus for RAG.
throughput at fleet scaleIn the same 30B test with requests running at once, the tier finished answers per second versus for RAG.
vs per 1k queriesAt that same -at-once point: take the box's hourly price and divide by the answers it sustained per hour (formula below).

Fleet-scale throughput

The multi-tenant regime, not a lab best case.

We ramp in-flight concurrency up to on one serving GPU tier and measure both arms on the same model and retrieval. The RAG arm runs in the condition a busy fleet actually creates: a per-request nonce defeats the engine's prefix cache — exactly what other tenants' traffic does under churn — so RAG re-reads its context every request while the resident-KV arm does not.

Setup

Qwen3-30B-A3B (FP8), tensor-parallel across a 4-GPU tier (4× L4, $/hr). Same model and same retrieved documents feed both arms.

Measured, server-side

Time-to-first-token is the real time to the first streamed token. Throughput = completed requests ÷ wall-clock at each concurrency level — queueing included.

Cost formula

GPU $/1k queries = instance $/hr ÷ (sustained queries/sec × 3600) × 1000. The real running price of the box, not a list rate.

ConcurrencySmart CAG q/sRAG q/sCAG TTFTRAG TTFT

RAG throughput barely moves as concurrency climbs — it is compute-bound re-reading context on every request — so the resident-KV lead widens with load. Full ramp and the raw run outputs are in the source-of-truth results log.

Answer Quality

Same accuracy, different efficiency

Query answer accuracy is a statistical tie with a standard RAG architecture. On a 1,000-document haystack, resident-KV serving matches an equally-retrieved RAG within noise, while re-processing far less text per question.

Same retriever, both sides

A 1,169-document QASPER haystack, 500 questions, on an 8B model. One shared hybrid retriever (keyword + dense + reranker) feeds both arms — so the comparison isolates the serving path, not retrieval quality.

The tie

A paired significance test put the gap between the two well inside statistical noise — on answer quality, they're not distinguishable.

Independent judge

The result is corroborated by a second, independent frontier judge — a different, more capable model than the one under test, so the verdict doesn't depend on a lenient self-grade.

Reproduce it on your stack.

We'll share the connector, the conformance suite, and the prefix-cache benchmark so your own engineers can run the head-to-head on your traffic shape.

Get in touch