The machine learning community relies heavily on double-blind peer review and open rebuttals to evaluate scientific contributions fairly. OpenReview was adopted by major venues like NeurIPS, ICLR, and ICML specifically to foster transparency — allowing authors, reviewers, and Area Chairs (ACs) to engage in evidence-based discussions.
However, transparency on paper does not guarantee procedural due process in practice. What follows is a documented case study from the NeurIPS 2026 review cycle illustrating a severe failure mode in conference governance: the manufacturing of post-rebuttal reviewer consensus by an Area Chair despite total reviewer non-engagement during the official discussion window, alongside arbitrary score manipulation.
The Timeline of Events
1. Initial Reviews & Requested Experiments
The initial reviews raised standard technical questions and requested additional evaluations, specifically asking for native-task empirical benchmarks against state-of-the-art baselines to support our claims. Initial scores ranged up to a 6.
2. The Author Rebuttal
Within the official rebuttal window, our team conducted the requested native-task evaluations. We uploaded comprehensive empirical tables demonstrating significant, order-of-magnitude performance gains over leading baselines across standard benchmark datasets, directly addressing every technical point raised.
3. The 11-Day Score Flip & The “Clerical Error” Excuse
During the post-rebuttal window, Reviewer A suddenly altered their score from a 6** down to a **4 without posting a single line of new technical feedback or critique to justify the downgrade. When we wrote about this to the area chair, the area chair queried the reviewer and learned that it was merely an un-evaluated “clerical error” made when entering the score into OpenReview. It took 11 days — and required direct intervention from the Area Chair on the very last day of the author-reviewer discussion window — for the reviewer to acknowledge the mistake, claiming it was merely an un-evaluated “clerical error” made when entering the score into OpenReview.
While the Area Chair formally promised in writing that “I do not rely on this review in my final recommendation,” the paper was left saddled with an uncorrected phantom downgrade throughout the entire active discussion period.
4. The Discussion Period (Total Silent Ghosting)
During the entire 8-day official author-reviewer discussion window, zero reviewers responded to the rebuttal. Not a single reviewer acknowledged receiving the new experimental data, posted a follow-up question, or offered a technical critique of the submitted benchmarks. The official discussion log shows 0 author-reviewer interactions after our response was posted.
5. The Final Meta-Review & Rejection
When final decisions were released, the Area Chair issued a rejection. To cover for discounting Reviewer A’s invalid score, the AC relied on ghost feedback from the remaining silent reviewers, writing in the meta-review:
“Reviewer [B] maintained that concerns regarding evaluation and baseline fairness were not sufficiently resolved. Reviewer [C] further noted that the rebuttal emphasized specific metrics without adequately addressing error magnitudes…”
The Procedural Breakdown
This outcome represents a fundamental breach of academic due process across multiple levels:
- Factual Misrepresentation of the Record: The official OpenReview audit log proves that Reviewers B and C postedzero comments during the discussion period. Asserting on the record that a reviewer “maintained” a position post-rebuttal when they ghosted the thread misrepresents reviewer activity and invents a non-existent public consensus.
- Denial of the Right to Respond: If a reviewer secretly submitted a brand-new technical objection to the AC in a private chatafter the public window closed, using that objection as the basis for rejection violates core conference rules. Authors cannot defend their work against ghost critiques that were never entered into the public record during the active window.
- Unvetted Score Manipulation & Administrative Inertia: Allowing a reviewer’s unexplained score drop to sit uncorrected on OpenReview for 11 full days — only to dismiss it as a “clerical error” after the rebuttal phase closed — severely disrupts the evaluation process. It leaves authors reacting to phantom score changes while forcing Area Chairs to write awkward meta-reviews that rely on ghost feedback from silent reviewers to justify a pre-determined rejection.
When an Area Chair discounts a flawed review only to replace it with un-rebutted, phantom feedback from silent reviewers, the open rebuttal system ceases to function as a scientific dialogue and becomes a procedural fiction.
Peer review under high submission volumes is undeniably difficult, and Area Chairs face immense pressure. However, substituting unsubstantiated private resistance for transparent, evidence-based deliberation undermines author trust in major CS venues.
If authors invest substantial compute and labor to run requested rebuttal experiments, conference mechanics must enforce a basic rule: a silent reviewer who fails to refute submitted empirical evidence during the discussion window cannot be cited as having valid, standing objections.
We share this case study not to litigate a single paper decision, but to urge Program Chairs and conference boards to implement stronger automated checks on OpenReview — ensuring that meta-review summaries accurately reflect the public discussion log and that due process is preserved for all authors.