Academic Conferences are Overloaded.
Peer Review is Broken.
What Can Computer Scientists Do About It?
Posted on September 18, 2026 by Rohan Padhye (@moarbugs on X)
Lately I keep running into the same complaint from a lot of people: everyone is exhausted by the state of conferences and peer review in computer science academia, especially in the fields where I publish: software engineering, programming languages, and systems. Submission volumes keep climbing (ASE alone nearly doubled, from 587 to 1,181 submissions between 2024 and 2025, with acceptance rates flat), and reviewer loads climb with them. The numbers are vastly more staggering in the AI research community. I've felt this exhaustion myself, as an attendee, a reviewer, and an author, and that's what pushed me to write this post.
For a rigorous, evidence-backed treatment of what's wrong with peer review in computer science, I recommend Nihar Shah's survey “An Overview of Challenges, Experiments, and Computational Solutions in Peer Review”. It already does the data-gathering better than I could, so what follows is just my own opinions on what's broken and what, if anything, we can do about it.
What are conferences actually trying to do?
I think academic conferences are being asked to serve too many functions at once, at least four of them:
- Quality gating. Peer review should assess whether a piece of work is technically sound, so flawed or incorrect results don't spread through the community as accepted fact.
- Feedback and mentoring. The review process should give authors feedback that helps them improve their paper, methodology, or writing, doubling as informal mentoring for junior researchers.
- Reward and credentialing. An accepted paper is a credential: beans for those who like to keep count in hiring, promotion, tenure cases, and PhD thesis defenses. Conferences also often give out "distinguished" or "best paper" awards. All these credentials help researchers secure grant funding and keep the cycle going.
- Allocation of talk slots. For conferences (unlike journals), the accept/reject decision also decides who gets to give a talk in person. It allocates a scarce resource — presentation slots and seminar-room space — among far more people who want one than there's room for.
To me, this overloading is the root of the problem, and it's a problem with conferences as an institution, not just with reviewing. A conference paper is expected to be a stamp of scientific soundness, a mentoring exercise, a promotion-and-tenure credential, and a ticket to give a talk, all bundled into the same peer-review process.
The vicious cycle this creates
This conflation cascades into a chain of problems, each one making the next worse:
- Authors are, and always have been, incentivized to publish more. That's not new, and it's not caused by AI: it follows from the reward function above, where more accepted papers means more credentialing for hiring, promotion, and tenure. What's changed is that AI tools now make it much easier to act on that incentive, automating experiments, evaluations, and even the writing itself, which has driven a real explosion in submissions. This is especially pronounced in software engineering, where it's easy to find research questions that sound valid and haven't technically been asked before, but can be answered by a pipeline that's increasingly automatable end to end.
- Scaling up reviewing to match has hurt reviewer-paper matching. The standard response to rising submissions is to grow program committees, often to several hundred members, and split reviewing into multiple cycles a year, the way ICSE now runs two submission cycles for its research track. But with a much larger pool of papers and reviewers, and a roughly fixed cap on how many papers any one reviewer can take, it's much harder to match papers to reviewers with the right expertise.
- Mass reviewing has flattened what kind of paper gets accepted. When a large pool of time-crunched, not-always-expert reviewers has to get through an equally large pile of papers, an easy, formulaic criterion becomes the path of least resistance: does the paper show a number going up against a well-understood baseline on a familiar benchmark? That's a checklist a reviewer can apply without domain expertise, so it becomes the de facto bar for acceptance, and papers get judged in an increasingly cookie-cutter, monotone way. This disincentivizes non-traditional contributions: experience reports, reflections, papers that argue for why a problem is interesting, or ones that try to spark discussion rather than close things out with a clean empirical win. I think this is especially true in software engineering, where the field ends up rewarding papers that fit this "number go up" formula over ones that offer a genuinely unique perspective or engage with a significant, messy problem industry is actually facing.
- Reviewer anonymity keeps the stakes of writing a bad review low. It's not that reviewers have no incentive to do a good job: most people want to be conscientious. But since a reviewer's name is never made public, nothing rewards a thorough review or penalizes a lazy one. Nobody's career rises or falls on the quality of a review they wrote. Shah's survey documents where this leads: a non-negligible fraction of reviewers are now suspected of submitting low-effort, LLM-generated reviews outright, to the point that researchers have had to design statistical watermarks in the paper's PDF just to catch it.
- Accept/reject can look like a coin flip: Thanks to expertise mismatches, time-crunched reviewers, and anonymity removing any accountability, the outcome of reviewing decisions appears to be increasingly random. Nihar Shah's survey backs this up with a striking experiment. Splitting a set of NeurIPS submissions between two independent program committees found that 57% of the papers accepted by one committee were rejected by the other. Repeating the experiment years later, with an order of magnitude more submissions, found the same inconsistency.
- The long resubmission treadmill. Because review cycles are long and outcomes are largely random, rejected authors revise and resubmit elsewhere, adding more volume to the next conference's pile in a feedback loop that keeps the whole system growing. It also stretches out the time between when a paper is first written and when it's finally accepted somewhere.
- Conference talks are often outdated, and to me this is the most damaging consequence of all. In fields like software engineering, ideas and results can go stale within months, as new models and industry practices race ahead of them. By the time a paper survives several rounds of rejection and revision, gets accepted, and waits another three to six months for the conference, the talk on stage often reflects work the authors did one or two years earlier. Attendees hoping to learn about the latest research end up hearing ideas that are already stale. When I attend conferences, I am more interested in the "hallway track" because I want to know what folks are doing this month, not last year.
So what can we do about it?
My proposal is to pull the four functions I described earlier apart, so each one can be optimized on its own terms instead of being crammed into one accept/reject decision:
- Rate-limit submissions per author. ICSE tried a cap of three submissions per author back in 2017, but that clearly isn't holding: at ICSE 2026, one person co-authored 37 papers in a single cycle, even though 95% of authors submitted three or fewer. ICLR recently capped every author at 20 submissions for its 2027 cycle after a 68% jump in submissions. I'd set the bar much lower, at two or three per author per cycle, since that already covers the vast majority of authors and only targets the extreme outliers.
- Publish all sound research. A lot of the pressure in this system comes from treating a single paper as both a scientific contribution and a competitive credential. Any methodologically sound, well-written piece of research should just get published, at close to 100% acceptance, the way arXiv already works, instead of competing for a scarce number of conference slots. Call this the research-article layer: its only job is to get sound science out to the field, with no cap on how much good work gets in. Who reviews all this for soundness? Authors can invite their own reviewers, peers with no conflict of interest, to review the work publicly and non-anonymously, backed by a default AI-generated review (something AAAI-26 already piloted). Anyone else in the community is also free to submit a voluntary review at any time. Authors respond to all of it in an open, conversational back-and-forth until the quality bar is cleared, typically in days or weeks rather than months. Then it's out.
- Introduce a centraized platform for self-edited research highlights. It's a short document, more like a blog post or a statement than a paper, where researchers write about whatever they consider their best work over a recent stretch and point back at the underlying research articles, so anyone reading it can also see the full reviews and discussion behind each one. It serves two purposes: reward evaluation, whether that's a PhD, promotion, or tenure committee, or a community-run "quarterly distinguished highlight" award, and it's the unit you submit to bid for a conference talk slot. Nobody has to produce one on any cadence: a PhD student might reasonably put one together just once a year. Rate-limiting it at one per quarter keeps the incentive on quality over frequency. This lines up with the CRA's own best practices memo on tenure and promotion, which recommends evaluating candidates on their three to five most important publications rather than a running paper count.
- Apply for a conference talk slot using research highlights. When you have a highlight, you can bid to present it as a conference talk, as long as it reflects research concluded in the last two or three months, so what's on stage stays current. The program committee isn't gating quality here, since the underlying article already cleared that bar; it's judging whether the content would interest the people in the room, a quick review rather than a scientific one. Because a highlight already shows a track record of solid, quality-gated work, the talk itself doesn't even have to be about that exact result: it could just as easily cover newer work in progress that hasn't cleared the quality gate yet, as long as the committee trusts it will. If a highlight doesn't get picked for a talk, there's no resubmission: you just bid your next quarter's highlight at a future venue, keeping things fresh. Nihar Shah makes a similar case in his position paper arguing that peer review should separate publication from presentation.
- Make peer review public and non-anonymous. The same position paper argues for non-anonymous review tracks, and I agree: if a reviewer's name were attached to their review, sloppy reviewing would cost them something. To protect junior reviewers from retaliation, reviews could be co-authored, so a junior researcher can put their name on critical feedback alongside a senior reviewer who backs it. Authors and reviewers can then argue things out in an open forum, closer to Hacker News than a review portal, with the wider community voting articles up or down through ORCID-verified logins to keep out bots.
I don't expect any of this to happen overnight, or at all. I'm happy to get feedback or continue constructive discussions via email or social media. Hopefully academia can adapt to the changing times one way or the other.