Data Ethics Club: The System That Decides What Science Gets Published Is Breaking Down#

Article summary#

Peer review is the foundation which all published scientific findings rest upon. Authors submit their papers to a journal, where an editor selects independent expert reviewers to scrutinise the work’s methods, claims, and contribution. Reviewers provide recommendations for whether the work should be published, with the editor making the final decision. It is an imperfect process but an essential part of producing quality science which is verifiable, reliable, and reproducible.

Despite its vital importance, the system of peer review is under strain and may be quietly unravelling. Submissions to top journals are surging, straining the pool of qualified reviewers. To meet demand, editors need to either recruit less experienced people or put more pressure on the reliable ones. Either way, the quality of review drops. In response, authors take more chances and submit to prestigious venues that otherwise might have been ruled out.

Structural mechanisms in the peer review process make it difficult to respond appropriately to the increased demand. Firstly, peer review is unpaid, meaning that increasing wages is not an avenue to meet demand. Introducing paid reviews may not have the desired effect – paying small amounts can result in decreased incentives to participate and attracting the wrong people. Secondly, journals do not share the workload between each other. When a paper is rejected from one journal, authors can just work their way through others, with each venue generating waves of re-review.

To move towards fixing the problems with peer review, the article argues that editors need to place more emphasis on desk rejection before papers reach reviewers, conserving reviewer labour for papers that genuinely deserve scrutiny. Placing career recognition on quality reviewing could also improve the process, however, would require the scientific community to work together to renegotiate what it values.

Discussion summary#

What does the value of publication and peer review mean to you? What would you like it to mean?#

Peer review is essential to good scholarly hygiene, with benefits to both authors and reviewers. As authors, we’ve found it can provide useful guidance and suggestions that improve the quality of our research. As reviewers, we have gained valuable learning. It’s useful for following a wider scope of work in our field as well as improving our critical thinking.

Whilst there is value in peer review, we acknowledged that it is not infallible. The way that individual reviewers evaluate a piece of work is based on their experience, interest, and unconscious biases. This has benefits, for example, as reviewers have a different perspective to authors which may enable them to notice things that the authors may have missed. However, it can also mean that sometimes reviews are misinformed or inaccurate. Comments can be unfocused or marginal if reviewers haven’t invested enough attention and biases can arise if reviewers tend to be only from certain backgrounds, such as by neglecting the global south input. Bad reviewing can prevent the process from improving research, and sometimes even make it worse.

Reviewer anonymity helps to mitigate bias yet can also be a double-edged sword. There are benefits to being open, such as supporting accountability, and some journals are more open than others. We’ve had good experiences where we have received editor approval to publish a pre-print of a submitted paper, followed by open peer review where the comments are open for anyone to read.

Issues with peer review impact beyond just the authors and reviewers themselves because there is a prevailing belief that if something is published, it must be true. When attempting to influence policy or public opinion, research is given a weightier value if it is published in a well-known venue like Nature or the British Medical Journal (BMJ). However, strain on peer review means that research can take a long time to go through the publication system and by the time it is published, it is too late for the work to inform decisions. This can have a negative impact in urgent areas such as climate change.

Problems with peer review can lead to good research taking too long to get into the public domain, but it can also mean that bad research slips through. The publication of poor-quality work is especially problematic at the moment as the end users of research are changing. Large language models (LLMs) are increasingly citing journal articles, which is leading to some works being viewed more often than they were before LLM adoption. If LLMs are citing poor quality research, bad science can spread.

Poor research communication catalyses misinformation spreading. Authority bias arises when people accept something that is difficult to understand because they assume that the authors know more than them. A similar bias also happens in reverse, where authors think that the public won’t understand the nuance of their work. To improve the accessibility of research, we thought that there is a good argument for plain English summaries, which we are increasingly seeing in trustworthy sites like Mental Elf.

Overvaluing publication has led to an atmosphere of publish or perish in academia, which drives prolific publication. Publication count is viewed as a highly important metric for job applications. Yet, this is straining the publication system, leading to desk rejections growing more common which can be discouraging and mean that the researcher misses out on the feedback that comes from peer review.

Lots of venues are now saying that they have too many papers submitted to them, which underlines a need to shift publishing incentives from quantity to quality. Cycles can emerge where the same peer reviewers are asked repeatedly, exhausting those reviewers as well as the editors who run out of alternative people to pass on to.

The emphasis on publication is problematic as academic value cannot be reduced to publication output. Whilst more senior academics are less affected and can be agnostic about publish or perish, at the bottom of the academic ladder there is a rat race for survival. The lack of consistent regulation in the publication system impacts junior researchers the most, which means that those who are impacted the most have the least decision-making power in changing the model.

Given the issues with peer review and the publication machine, we wondered what the real value of human review is and how we can make the most of it. Maintaining review quality is key, but it is not very clear who reviews the reviewer. Editors are responsible to some extent, however, the role is not universally defined. Often, they cover the gap between the authors and the reviewers.

What do you think of the contrast between how editors are described as not being paid enough with the daycare study that suggests paying reviewers wouldn’t be effective?#

To get people to recognise the importance of their role as a reviewer, we need to cultivate a system that celebrates it more. Reviewer names could be added to papers, or prizes can be given to good reviewers such as free conference tickets, which is done at some computer science conferences. Some journals have public reviews or reviewer credits, for example, the Journal of Medical Internet Research (JMIR) have “karma” points, which incentivise reviewers by boosting their reputation. Peer review tends not to be highly valued on a CV, but it should be made clear in research job descriptions that it is a key part of your role.

Cash in hand for reviewers might be effective; reviewers do have bills to pay and should be valued by society for the work they do. There is profit being made from publication, by journals and other actors like AI companies who use published works for training models. It is reasonable to expect to be paid for doing a professional task. Arguably, not paying workers for something that is being profited from is an exploitation of labour.

However, it is not necessarily true that financial compensation is required. Some trials of alternative peer review models have solved one problem only for others to crop up. From an ethical perspective, there is an element of altruism involved in reviewing. Issuing payment for reviewers can attract people who are purely financially motivated and churn out poor quality reviews for the money only.

Whether people ask for payment for reviewing might be up to the individual. Yet, we wondered whether it is fair for the individual to shoulder that burden. As the primary financial beneficiaries, perhaps the onus should be with the journal or company publishing the research. Not paying reviewers or leaving it up to individuals whether they decide to ask for compensation, is self-serving for journals as they save costs by doing so.

As Gross put it, many of the interventions “would require coordination among large groups that typically don’t coordinate.”’ Do you agree with this? Do you think a coordinated solution is possible?#

Massive coordination is required to change the publication system. Even if everybody agrees that change is needed, the system is highly entrenched and involves processes which are hundreds of years old.

Alternative models to the traditional peer review process have been tried. The Cochrane Collaboration has a rigorous and systematic review process, where reviews are updated over time to include new evidence and methods, and reviewers work with authors to develop the research. People still will want numbers that can go up, so maybe something like aggregate index of rigour could be used.

It is crucial that peer review scrutinises methods. Very often, the methods in a paper are underemphasised. However, if the methods are wrong or flawed, the whole paper becomes invalid. Paperstars evaluates scientific papers through methodological soundness and transparency instead of citation counts. It provides a platform for researchers to submit anonymous ratings and reviews of papers.

Mechanisms for getting research into the public domain faster are also employed. In rapidly advancing domains like machine learning, it is common for people to publish pre-prints. World Weather Attribution conducts rapid attribution studies and makes results public as soon as they are available to inform discussions about climate change and extreme weather.

Widening the net of people involved in reviewing could help filter out the volume of works submitted for publication and improve the quality of review. There are opportunities to make use of junior academics and students, as they will have different perspectives to researchers that are more established and thereby entrenched in their ways. There is a lot of value in including new academics in peer review, but we’ve found that their perspective is often overlooked. Involving post doc students could be a way to include inputs from a wider variety of backgrounds, like people from the global south. We had differing views on whether junior or senior researchers would be more likely to provide detailed reviews. Newer researchers might take the role of reviewer more seriously and dedicate more energy to thinking through their assessment. More experienced researchers could be better placed to admit when they didn’t understand something. However, even delegating specific reviewing tasks to juniors can be useful, such as asking them to check for plagiarism or confirm citations. This would free up senior reviewers and editors to focus on in depth analysis of legitimate works

Venues are increasingly using LLMs to generate reviews. LLMs could be a useful avenue to outsource review work, easing the burden on researchers and providing additional viewpoints. However, in our experience LLM reviews have not been as useful as those from human reviewers. We’ve found that LLMs can pick out mediocre points, but are not very good at being novel, resulting in reviews that are more nit-picking than constructive. LLMs can be very sycophantic, and it is difficult to get nuanced outputs even if you tailor your prompts for this. Hallucinations are also still an issue.

Other ways to reduce the burden on peer review include checks and submission restrictions. Venues sometimes have registration quality checks, but it often isn’t clear what those checks consist of. Some grants have systems where applicants can only be rejected so many times before they are barred from submitting for an amount of time. arXiv now has a one-year ban on submissions from authors that don’t properly check AI-generated content in their work. For computer science papers, having a Git repository is often a good sign that the work is legitimate. Pre-prints can be a useful accountability tool as the version can be updated with each submission, creating a change history of the paper.

Attendees#

  • Huw Day, Postdoc in Digital Health, University of Bristol: LinkedIn, BlueSky

  • Jessica Woodgate, PhD Student, University of Bristol

  • Amy Joint, Clinical Trials publishing manager, Springer Nature: LinkedIn

  • Danny Schnitzler, postdoc in Stefan Lab (Medical School Berlin) and Paperstars

  • Melanie Stefan, Computational Neurobiologist, Medical School Berlin

  • Kamilla Wells, Citizen Developer / AI Product Manager, Brisbane

  • Naomi Cornish, clinical academic PhD student at University of Bristol

  • Patricia Rubisch, Computational Neuroscience postdoc in Stefan Lab (Medical School Berlin)

  • ZoĂ« Turner, Senior Data Scientist (NHS)

  • Nabeel Akram, PhD student in Aerosol Science, University of Manchester

  • Zosia Beckles, Research Information Analyst, University of Bristol

  • Paul Matthews, Lecturer in Data Science, UWE Bristol