Can national research assessment systems survive the AI storm?
Different types of national assessment systems face different threat levels
By Alex Rushforth, Paul Wouters, and Ludo Waltman (RoRI and the Centre for Science and Technology Studies, Leiden University)
In March 2026, a new edition of the Strategy Evaluation Protocol (SEP) was introduced in the Netherlands. In this post we ask what AI (large language models and Agentic AI) means for the future of national research assessment systems: the large-scale, typically state-organised periodic assessments of the quality and impact of research produced by universities and public research institutions, often linked to funding allocation and accountability. When the data these systems rely on is corrupted, and the evaluation processes they use are being re-shaped by AI, the foundations of these systems are at risk. Different types of national assessment systems face different threat levels. Some systems, such as the Dutch SEP, seem more or less future-ready, but many others are in need of fundamental reform.
The key pillars of knowledge production and governance are in trouble. Evaluation, funding, and scholarly communication are each facing serious challenges to their legitimacy, and their challenges are compounding one another. Take the paper mills operating at industrial scale that are currently flooding the literature with seemingly plausible looking but low-quality papers. The scale is startling: the US public health database NHANES (National Health and Nutrition Examination Survey), for instance, saw submissions using its data surge from 4 per year (2014-2021) to 190 in the first ten months of 2024. A recent study reported that the numbers of retractions have remained alarmingly high in recent years, with the growth rate of retracted papers significantly outpacing that of overall global publications. According to the study large language models (LLMs) are accelerating this dynamic: between 2024 and 2025, over 2,100 retractions were associated with AI-generated content, with retraction notices citing “tortured phrases”, “nonstandard phrases”, and text “generated by a large language model”.
Trusted platforms are beginning to collapse under pressure. The open access publishing platform Hindawi, acquired by Wiley in 2021, had to retract more than 11,000 papers and close down 19 open access journals before shutting down its platform entirely. This case illustrates the pressure on an already overburdened peer review system, in which the gap between claimed principles and actual practice is becoming increasingly hard to ignore. AI is widening that gap. The 2026 International Conference on Learning Representations (ICLR) permitted authors and reviewers to use AI and found it was confronted with 21% of its 76,000 reviews being written solely by AI agents, and more than half with partial AI generated text, leading to a deluge of verbose and nonsensical reviews.
Combined with the perverse effects of quantitative performance indicators and the hypercompetition fuelled by rankings and “publish or perish” cultures, these pressures converge on national research assessment exercises with particular force. These exercises often sit at the intersection of scholarly communication, evaluation, and funding distribution1, which are all, in different ways, in crisis. Important reform movements calling for less reliance on metrics and greater use of peer review do however pose some risks of increasing pressure on an already overburdened system, while the legitimacy of that system, long sustained by an idealised account of what peer review delivers, grows harder to defend. LLMs have landed on ground that was already shifting, amplifying existing strains and making the case for deferral or selective attention increasingly untenable. Not all types of national assessment systems are equally exposed, however, and we argue that exercises geared more towards strategic intelligence, advice, and dialogue are better equipped, though not wholly immune, to cope with these threats.
How AI threatens national research assessment systems
Currently, national research assessment exercises are exposed to AI-related pressures in two ways.
The first is pollution of the research outputs many of these systems rely on. AI techniques such as LLMs and Agentic AI make it possible to generate large volumes of plausible-looking research outputs at scale. This already seems to be under way: a recent blogpost reported the number of articles submitted to journals using the ScholarOne submission system rose by 33% between 2025 and 2026. Editorial processes are unlikely to cope with this growth, meaning AI-generated articles will increasingly enter the published literature. For systems that use publication counts or citation metrics as proxies for research quality, this strikes at the validity of the data on which assessments of the merits of traditional research outputs and contributions depend.
The second is infiltration of the evaluation processes themselves. LLMs are in all likelihood already being used to support the writing of submissions and perhaps in some settings the conduct of reviews (in other closed peer review settings the evidence is rising). This is happening largely in the shadows. This raises questions about what exactly is being evaluated when AI mediates both the production and evaluation of knowledge. Peer review has always rested on an idealised account of what it delivers, even if the evidence has long pointed to a more complicated reality. AI sharpens this legitimation problem, with the black-boxed nature of LLM algorithms sitting awkwardly alongside perennial calls for transparency and calling for reassessment of the primary status of standalone human expert judgment.
Assessing threat levels: not all systems are equally exposed
Let us now consider how these threats might bear down on different types of national assessment system. The RoRI AGORRA project suggests there are eight dimensions along which national assessment systems vary and that differences are considerable. For present purposes, however, we focus on three major ideal types of system and construct scenarios that might plausibly befall them.
Red alert: quantitative, output-based systems
Systems that link funding allocation directly to publication output in indexed databases or journal lists are the most immediately exposed. As the literature these systems measure becomes increasingly contaminated by AI-generated content, the key parameter used for resource distribution is fatally undermined. An “arms race” between generation and detection ensues - one that does not favour the latter, and no satisfactory technical fix is found.
Significant risk: large-scale peer review exercises
Peer-review-driven accountability models like the UK’s REF and Italy’s VQR face a serious but different risk exposure. AI-assisted writing of submissions, and quite possibly of reviews themselves, begins to happen largely in the shadows: acknowledging their use by reviewers openly risks litigation over decisions mediated by black-boxed proprietary algorithms. A likely response is formal prohibition of the tools by peers, producing patchy and unverifiable enforcement, with management of the problem a ‘dirty secret’ rather than a governance challenge.
Some safeguards remain - expert panels have discretion to discriminate high quality work from the AI research “slop” (some of which may also be screened out by the submitting institutions).
Meanwhile, existential questions about what evaluation actually consists of become harder to avoid: if machines can correlate reasonably well with expert assessments, questions get asked about what exactly justifies the cost and burden of large-scale expert review. Cases are put that these tools generate superficially persuasive evaluation signals, shorn of core attributes that would make it genuine evaluation. At the same time there is openness to experimentation. However, while parts of the sector try to speak the language of augmentation (exploring whether LLMs may support improvements in expert-led evaluation in circumscribed ways), the governmental bodies that authorise and resource these exercises (tired of costly, burdensome exercises that empower academic elites) are persuaded by the speedy promises of LLMs and enhanced analytics to end large-scale peer review, putting the large and costly national exercises on life-support.
Lower risk: formative advisory exercises
Systems oriented toward strategic advice and organisational learning rather than competitive ratings and funding verdicts have certain structural buffers that the two aforementioned assessment models lack. An exercise like the Dutch Strategy Evaluation Protocol (SEP) could in principle see department staff connive to outsource their evaluation to AI, but mitigations like institution-wide self-evaluation reports and face-to-face dialogue between expert panels and evaluated units create social accountability that makes this very unlikely - checks and balances that purely algorithmic or behind-closed-doors processes (separating the evaluator from the evaluated) lack. Because results are not directly tied to funding or reputational ranking, the political stakes of AI involvement are lower, and there is more room for locally negotiated, transparent uses of AI that can be discussed openly rather than driven underground.
The need for fundamental reform
AI is not the cause of the challenges facing national research assessment, but it is making them impossible to ignore any longer. The layering of LLMs and Agentic AIs onto systems already under pressure is forcing a reckoning with structural problems that have accumulated over many years.
No one model of national research assessment system is completely immune to these pressures. But the threat levels differ considerably, and these differences shape how urgently reform is needed and what forms it should take. Systems that link funding tightly to outputs measured through increasingly compromised data, or that depend on expert judgment processes they cannot be transparent about, face a credibility problem that incremental reform is unlikely to solve. For these systems, a serious exploration of different evaluation and funding models is urgently needed.
Even the more formative systems, despite their structural buffers, cannot afford complacency. The Dutch SEP and exercises like it need to maintain a sceptical and reflexive focus on what AI means for their own processes and the risks it introduces, however differently those risks are configured compared with more summative, resource distribution-focused national assessment models.
Moving toward systems that generate strategic intelligence rather than competitive verdicts - whether formative for the evaluating agency, as Australian Research Council's proposal for a Research Insights Capability envisages, or formative for the evaluated institution, as the Dutch SEP demonstrates, are directions worth taking seriously.
The challenges we present will not resolve themselves, but will force choices that previously may have seemed deferrable. Whether national research assessment systems confront those choices before cascading pressures become irreversible is a question that should not be put off any longer.
Claude Sonnet 4.6 was used to assist with drafting, editing, and restructuring the blog text and as a search engine to help identify concrete instances of AI problems. The authors take full responsibility for all evidence, arguments, and conclusions that are presented.
Pressures wrought by AI are also being reported in the context of grant funding instruments.


