arXiv:2606.21807, cs.CL. 176,864-row reader-by-compression matrix released

Fixed RAG Compression Collapses Measured Reader Scaling

Sugam Panthi and Rabab Abdelfattah

Fixed-Artifact Reader Replay: byte-identical compressed evidence replayed across 8 to 20 reader panels

QR code: Scan for paper

Scan for paper

The problem

Fixed compression shrinks the measured gap between readers

Compression is evaluated as a deployment layer: does a shorter context preserve accuracy? The same layer is then reused while readers are compared. These roles conflict. A compressor can raise several pipeline scores while absorbing the reader upgrade the evaluation was built to measure.

RAW EVIDENCE 12.6 44.4 31.8pp SAME PAIR, FIXED RECOMP 36.4 44.2 7.8pp

7.8pp is all that remains of a 31.8pp raw reader contrast after one useful in-domain RECOMP artifact. Same 500 HotpotQA rows, same candidate pools, same endpoint pair.

Method

Fixed-Artifact Reader Replay: change only the evidence policy

saved candidate pools per row rows, retrieval, prompts, and scorer fixed before replay raw evidence the declared reference policy one compressed artifact per row content hash must match for every reader record replay both policies across the panel 8 to 20 readers per panel, 11 model families upgrade retention ρ = Δcomp / Δraw excess ranking flips vs raw replay floor

The endpoint pair is selected by raw score only, before compressed scores are inspected, and its compressed span reuses the same pair. Excess flips subtract the reversal rate of repeated raw-evidence evaluation on held-out rows, so ordering changes are counted beyond the panel's own row-sampling floor.

The primary study covers 4 benchmarks and 5 compression families in a released 176,864-row interaction matrix. A method enters a panel only when every reader record has the same content hash for that row-level artifact. The analysis code performs this check; a shared method label alone does not admit the panel.

Core result: prespecified confirmatory panels

Compression raises low-reader scores more and reduces the reader gap

HOTPOTQA, FIXED RECOMP, EM 0% 0% 25% 25% 50% 50% identity retains 24.5% of a 31.8pp contrast 20 readers, same rows, same pools MUSIQUE, SHARED SUMMARY, F1 0% 0% 25% 25% 50% 50% identity retains 0.8% of an 18.7pp contrast 15 readers, same rows, same pools compressed score raw score
Each point is one reader under raw and byte-identical compressed evidence; ringed points are the raw-selected endpoints. HotpotQA retention is 0.245 for EM and F1 (family-by-row intervals [-0.007, 0.379] and [0.050, 0.344]); the paired EIV correlation between raw score and compression gain is -.961 (EM). MuSiQue F1 retention is .008 [-.117, .270]; EM agrees in direction but is imprecise. The panel mean rises in both cases: this is floor-lifting with reader-dependent gains, not a common failure floor. Data: paper Figure 2 and Table 2.

Decision consequence: 5,000 row splits

Compression reverses reader orderings beyond row noise

0 +10 +20 +30 +40 +50 raw replay floor HotpotQA, EM +16.7 HotpotQA, F1 +16.7 MuSiQue, EM +32.7 MuSiQue, F1 +21.2
Median excess flip rates across 5,000 disjoint row splits, for reader pairs separated by at least 5pp; bars are 95% panel-conditional split-stability intervals. Raw replay flips only 1.5% of eligible HotpotQA pairs, against roughly 18% under compression; the MuSiQue raw floor is zero. Data: paper Figure 3.

One concrete decision change (MuSiQue F1, raw top-3 shortlist):

DeepSeek R1 Distill Llama 70B wins under raw evidence. GPT-4.1-mini wins under the shared summary.

The account: row transitions

Rescue and damage coexist inside a rising average

HotpotQA 19.1% rescued: raw-wrong, correct after compression (share of all row-reader pairs) 35.7% damaged: raw-correct pairs that become incorrect MuSiQue 22.9% rescued: raw-wrong, correct after compression (share of all row-reader pairs) 33.8% damaged: raw-correct pairs that become incorrect
The two rates have different denominators and do not net into an accuracy change; they show substantial opposing movement inside the same average. Token F1 mass agrees: HotpotQA moves 22.0% positive against 12.6% negative, MuSiQue 32.3% against 3.4%. Different readers reach similar compressed averages through different paths. Data: paper Table 3.

Validity check: four HotpotQA policies

Not every compressor collapses the curve

EXIT .836 RECOMP extractive .245 LLM summary .181 RECOMP abstractive .138 endpoint EM retention (1.0 = upgrade fully preserved)
Endpoint retention spans .098 to .836 across four HotpotQA evidence policies and two metrics; EXIT preserves most of the curve and its excess-flip intervals cross zero. The later 13-reader TriviaQA panel retains .428 (EM) and .391 (F1). The claim is that fixed compression can distort a comparison, not that every compressor does, so each deployed compressor needs its own audit. Data: paper §6.

Takeaway

Audit fidelity separately from utility

  1. Freeze the evidence. One retrieval pool and one compressed artifact per row, with recorded hashes; regenerated summaries are different evidence.
  2. Select weak, middle, and strong readers by raw score, before inspecting compressed scores.
  3. Replay both policies and report pipeline means, upgrade retention, excess flips, and rescue and damage.
  4. A rising average is not evidence of a preserved comparison: RECOMP helps several pipelines while hiding most of the HotpotQA upgrade.
  5. Raw evidence is the declared reference policy, not a claim about intrinsic model capability.

TL;DR

What we did

Fixed-Artifact Reader Replay: freeze rows, retrieval, and one byte-identical compressed artifact per row, then replay raw and compressed evidence across 8 to 20 readers. Four benchmarks, five compression families, and a released 176,864-row interaction matrix.

What we found

One useful in-domain RECOMP artifact cuts a 31.8pp HotpotQA reader contrast to 7.8pp (24.5% retention); MuSiQue F1 retains 0.8%. Both benchmarks reverse separated reader orderings beyond the raw replay floor, up to +32.7pp excess flips.

What you should do

Never compare readers behind an unaudited compression layer. Freeze the artifact, report raw and compressed scaling, and quantify retention and excess flips; a compressor that raises the average can still absorb the very upgrade you are measuring.

+ and − zoom, 0 fits, scroll pans, ⌘/ctrl + wheel zooms, Esc or X closes