None of ARR’s 69,781 Published Paper Scores Is a 5.0

Type ResearchState cooking

No ACL Rolling Review paper has ever scored a 5.0. Across five years and 69,781 scored submissions, the top of the scale is empty. That does not mean reviewers hate everything. A 5.0 here is an average, and reaching it would take a paper whose whole review panel converged on a perfect score, which in five years never happened.

It is still worth sitting with, because on a different scale it does happen. ICLR publishes every reviewer's individual number instead of averaging them, and in 2025 the relighting model IC-Light, by ControlNet's creator, drew a straight 10, 10, 10, 10 from all four reviewers. ARR's averaging into a single 1-to-5 number makes the same outcome almost arithmetically invisible.

I found the empty ceiling while checking whether my own middling scores were harsh or normal. The important detail was how ARR constructs the published number: it averages a paper’s reviews before adding that paper to the histogram.

The data is public. ACL Rolling Review runs a stats dashboard backed by the open acl-org/arr-health repository, which publishes a per-cycle score histogram at iterations/<year>/<month>/review_scores.csv. Every number here is reproducible with a loop over those files.

These are paper scores, not reviewer scores

The number in the ARR histogram is one aggregate score per paper, not one per reviewer. This is the thing I nearly got wrong, and it changes what every chart below means. The histogram's total for each cycle equals the count of active submissions, not the count of reviews, and the real number of individual reviews is three to six times larger.

Cyclehistogram totalactive submissions (papers)actual individual reviews
2026 May13,66813,66842,250
2025 Feb7,3217,32123,126
2024 Jun4,7744,77415,234

The match to the paper count is exact, every cycle. So each row is one number per submission, the aggregate of that paper's three-or-so reviews (which is why the values still land on clean half-point steps). The 69,781 figure is submissions summed across 35 cycles, not reviews, and a resubmitted paper counts once per cycle. The individual reviewer scores that go into each average are not published anywhere.

That distinction matters because it kills the tempting headline. "No reviewer gives a five" is unsupported: individual 5.0s almost certainly exist and get averaged away. What the data actually shows is that no paper ever earns a 5.0 average, which is a claim about consensus, not about reviewer generosity.

No published paper-level average reaches 5.0

Across five years, the top of the aggregate scale is unused. Of 69,781 scored submissions, exactly zero reached a 5.0, and only 28 (0.04%) reached a 4.5. Anything at or above 4.0 is 1.4% of all papers. The nominal 1-to-5 scale behaves as a 2-to-3 scale: 83% of every scored submission lands at 2.0, 2.5, or 3.0, and the single most common score is 2.5.

ARR: aggregate review score per paper, 2021–2026
No ACL paper ever averages a perfect score: zero 5.0s in 69,781 submissions.
0.56%1.04.3%1.519.5%2.033.7%2.529.7%3.010.5%3.51.3%4.00.04%4.505.0no paper 5.0
Distribution of per-paper aggregate review scores (one score per submission, the average of its roughly three reviews) across all 35 ACL Rolling Review cycles, May 2021 to May 2026 (69,781 scored submissions, public ARR dashboard). The 1–5 scale behaves as a 2-to-3 scale: 83% of papers land at 2.0–3.0, the mode is 2.5, and the 5.0 bin is empty. A 4.5 appears 28 times (0.04%). These are paper aggregates; individual reviewer scores are not published.

Part of this is real and part is arithmetic. Averaging three scores mechanically pulls a paper toward the middle: to average a 5.0, every reviewer has to give roughly a 5.0, so the extremes wash out before they reach the histogram. The empty ceiling is therefore weak evidence about how any individual reviewer behaves and strong evidence about how rarely three reviewers agree a paper is flawless. Read it as a statement about consensus. Nothing in five years of ACL submissions cleared the bar of unanimous enthusiasm.

The pattern is not an artifact of tiny early cycles. It holds where it carries weight: the November 2021 cohort (2,585 papers), January 2026 (9,177), and May 2026 (13,668) all have zero 5.0s.

Single meta scores have wider tails than averaged review scores

Each paper carries two numbers: the aggregate of its reviews, and a single meta score from the area chair. The meta score reaches the top of the scale far more often. A paper's averaged reviews land at 4.0 or above 0.8% of the time; its meta score does so 7.5% of the time.

ARR: averaged reviews vs meta, per paper, 2025–2026
A paper's meta score swings wider than its averaged reviews.
avg reviewmeta (AC)1.01.52.02.53.03.54.04.55.036%22%fence: −14 pts7.5%0.8%champion: ×9.5
Per-paper aggregate review score vs the area chair's meta score, half-point-scale era only (ARR switched the meta scale from integer to half-points in Feb 2025), 43,458 papers. The meta score reaches 4.0 seven times more often (7.5% vs 0.8%) and cuts the hedged 2.5 fence from 36% to 22%. Caveat: the review number is an average of three scores and the meta is a single score, so a single score's higher variance explains part of the wider tails, not area-chair generosity alone.

Before reading this as "area chairs are the generous ones," note the confound I could not remove. The review number is an average of three scores; the meta number is one person's single score. A single score has more variance than an average of three by construction, so the meta distribution has fatter tails on both ends: it reaches 4.0 more often, and it also drops to 2.0 or below more often (26% versus 23%). Some unknown share of the gap is that arithmetic, not area-chair courage.

What survives the confound is the middle. Averaged reviews pile onto the 2.5 fence: 36% of papers sit exactly there. Meta scores refuse it, dropping the 2.5 share to 22% and pushing that mass outward. The area chair is the one point in the pipeline that resolves a hedged 2.5 into a verdict. I restricted this comparison to the half-point-scale era, February 2025 through May 2026 (43,458 papers), because ARR used an integer-only meta scale before then and mixing the two manufactures a difference that is really a scale change.

Mean paper scores stayed between 2.24 and 2.80 as volume grew 100-fold

The one finding here that needs no caveat is stability. A single cycle grew from 23 scored papers in mid-2021 to 13,668 in May 2026, more than a hundredfold, and the mean score did not follow. It stayed inside a narrow 2.24-to-2.80 band the entire time. The two lowest points, October 2023 and July 2025, are small off-cadence cycles, not a trend.

ARR: mean review score vs volume, 2021–2026
ARR grew more than 100× in five years. The average score didn't move.
0k5k10k2.02.52.9202120222023202420252026mean holds near 2.613,66823
Per-cycle scored-paper count (bars, left axis) and mean per-paper aggregate review score (line, right axis) across 35 ACL Rolling Review cycles, May 2021 to May 2026. Volume rose from 23 scored papers in mid-2021 to 13,668 in May 2026 while the mean stayed inside a 2.24–2.80 band. The two lowest means (2023/10, 2025/07) are small off-cadence cycles.

This is the observation David Jurgens flagged from the same public data. I read the flatness less as reassurance than as a property of a compressed instrument. When paper-level scores are averages clustered on a two-point band around "revisions needed," the mean has almost nowhere to move regardless of what gets submitted. A flat average under a 100x load increase is what a compressed scale looks like, not proof that quality held.

What this data cannot tell you

These are aggregated per-paper histograms, and the aggregation is the whole caveat. There are no individual reviewer scores, no paper identifiers, no accept or reject labels, no review text. You cannot recover how any single reviewer scores, measure disagreement inside a paper's review set, or join a score to an outcome.

It also cannot answer whether an automated reviewer would reproduce this distribution. There are no matched papers to score, and ACL reviews are public on OpenReview, so a model may have seen a paper's reception during training. That is contamination, not calibration, and it is the reason I keep my own reviewer-model work on held-out material rather than published venues. The honest scope of this post is one sentence: this is the distribution of paper-level aggregate scores across five years, and no paper ever averaged a perfect one.

Sources

Thanks for reading. If you enjoyed this piece, consider sharing it.