QED Science Claims It Can Rank the Top 1% of Preprints. Should Researchers Trust It?
QED Science ranked 57,455 life-science preprints and named 574 to its top 1%, despite unresolved questions about how its AI defines scientific quality.
The Tel Aviv startup says its system evaluates originality and validity without seeing authors, affiliations, citation counts, or journal names. That promise puts the QED Score against the prestige signals that researchers already use, often reluctantly, to navigate an overwhelming literature.
The central question raised by the Horizon Nature coverage is not whether artificial intelligence can read papers faster than people. It clearly can. The question is whether a proprietary score can become a trustworthy filter before independent researchers fully test its assumptions.
QED Science turned a year of preprints into one ranked list
QED Science has converted scientific triage from a reading problem into a ranking problem.
The company launched “The 1%” on June 24, 2026. It describes the project as a selection of the most original and scientifically valid life-science preprints posted during a 12-month period.
QED analyzed 57,455 manuscripts submitted to bioRxiv between May 2025 and April 2026. The final collection contains 574 papers, matching the top 1% of that scored population.
A preprint is a manuscript released before formal journal peer review. It can circulate results quickly, but readers must assess the evidence without the assurance associated with editorial review.
That speed creates both value and risk. Researchers can encounter important findings months before journal publication, yet they also face more material than any person can evaluate closely.
QED says a conventional review of its full corpus would require between 860,000 and 1,030,000 hours of expert work. That estimate equals more than 400 researcher-years, according to the company’s validation paper.
Its system completed the assessments in hours. That difference explains the appeal, especially for funders, research teams, and companies monitoring fast-moving fields.
The result is not an acceptance decision. QED describes The 1% as an early signal for deciding which manuscripts deserve attention.
That distinction matters. A journal editor can request revisions, consult reviewers, inspect disclosures, and reject a paper after finding a fatal problem. A score condenses a different process into one number.
QED’s public list spans bioRxiv’s life-science categories. Neuroscience contributes 83 selected papers from 11,058 evaluated manuscripts, while cell biology contributes 55 from 3,408.
Other categories include genomics, immunology, microbiology, cancer biology, bioinformatics, genetics, and synthetic biology. The distribution reflects both field size and QED’s assessments.
QED also reports adoption beyond the ranking project. As of May 31, its platform served more than 10,000 laboratories across over 1,500 institutions in more than 70 countries.
Those figures come from QED, not an independent usage audit. They nevertheless show that AI-assisted manuscript assessment is moving beyond a laboratory experiment.
The Horizon Nature question therefore reaches beyond one list. If researchers use QED to choose what they read, cite, fund, or develop, its judgments can shape attention before peer review begins.
That is the real change. QED has not merely summarized preprints. It has proposed an automated gate at the entrance to scientific attention.
Why the Horizon Nature debate matters now
The pressure comes from a widening gap between how quickly research appears and how slowly experts can evaluate it.
Preprint servers let scientists share findings without waiting for a journal’s full review cycle. That model became especially visible when researchers needed rapid access to emerging biomedical evidence.
It also removed a familiar sorting mechanism. A newly posted bioRxiv manuscript has no final journal placement, established citation record, or completed peer-review history.
Researchers compensate with imperfect signals. They recognize laboratories, institutions, authors, methods, and subject areas. They follow social recommendations and inspect whatever rises through professional networks.
Those habits can save time, but they can also reproduce prestige bias. Work from a famous institution receives attention more easily than equally strong work from an unfamiliar group.
QED is targeting that weakness. Its launch announcement says every manuscript is anonymized before scoring.
The system strips author names, affiliations, and other identifying information. It also ignores citation counts, publication venues, and journal reputation.
Oded Rechavi, QED’s co-founder and lead scientist, argues that scientific quality should be judged through the work itself. Blind evaluation is meant to prevent reputation from standing in for evidence.
That goal has a credible foundation. Human review can be influenced by institutional status, professional relationships, language, and expectations about who produces important work.
However, anonymity does not automatically create neutrality. A manuscript can contain indirect signals about its origin, including research topics, equipment, datasets, writing patterns, and references.
An AI model can also inherit associations from its training data. Removing the institution’s name from the input does not erase every relationship learned during model development.
Previous research has shown why that distinction matters. A 2024 study tested GPT-3.5 on 30 medical abstracts paired with different affiliations.
The researchers generated 232,500 individual reviews across randomized affiliation conditions. Their affiliation-bias study found that institutional information affected the model’s judgments.
QED’s anonymization addresses the most direct version of that problem. Yet the company has not publicly exposed enough system detail for outsiders to map every remaining pathway.
The scale problem is also worsening. Generative AI has reduced the effort required to draft technical text, create plausible citations, and produce polished manuscripts.
That does not mean every AI-assisted paper is unreliable. It means surface fluency has become an even weaker signal of scientific quality.
Richard Sever, a co-founder of bioRxiv, described a darker possibility in reporting about synthetic scientific content. Preprint systems become harder to use when manufactured papers overwhelm genuine findings.
Automated evaluation appears almost inevitable in that environment. Editors and scientists cannot answer unlimited machine-generated submissions with unlimited human review.
QED is offering one version of the response: use AI to decompose every manuscript, test its claims, and direct scarce human attention toward the strongest candidates.
The pressure falls first on researchers who conduct literature reviews. It also reaches journal editors, funders, biotechnology teams, and anyone making decisions from early evidence.
These groups need faster filtering. They also need to know why a paper passed the filter, what the system missed, and how often its errors repeat.
A ranking can reduce information overload while creating a new concentration of influence. If one score becomes widely trusted, mistakes at the scoring layer can redirect attention across an entire field.
That is why the debate has arrived now. Scientific publishing needs machine-scale triage, but the standards for validating that triage remain unsettled.
QED Score challenges prestige, not peer review
The strongest case for QED is not that it replaces peer review, but that it tests papers more directly than journal rank does.
QED Score evaluates two stated dimensions. Originality measures how far a finding advances existing knowledge, while validity measures whether the evidence supports the manuscript’s conclusions.
The company says its multi-agent architecture first decomposes each paper into a meta-claim, main claims, related claims, and supporting experiments.
Specialized AI agents then examine different features in parallel. They look for figure inconsistencies, statistical weaknesses, conflicts with existing literature, alternative hypotheses, and reporting-standard problems.
A verification layer rechecks flagged issues. A scoring layer grades individual claims, and an aggregator combines those assessments into a calibrated percentile.
A score of 80 means the manuscript ranked above 80% of the reference corpus. It does not mean the paper has an 80% probability of being correct.
That distinction is easy to lose when a complex judgment becomes a familiar number. Percentiles describe relative position, not absolute truth.
QED tested its metric in three studies. The first used 925 published papers from 185 authors that domain experts had placed into Limited, Satisfactory, or Strong categories.
The scoring pipeline did not receive those labels. QED reports an area under the curve of 0.867 when separating Limited papers from the other categories.
Area under the curve, or AUC, measures how well a system distinguishes two classes. A value of 0.5 represents chance, while 1.0 represents perfect separation.
For 795 papers with both QED Scores and journal-rank data, QED reported stronger results on two comparisons. It scored 0.863 against 0.804 for identifying Limited papers.
For Strong papers, the results were much closer. QED reported 0.782, compared with 0.774 for the journal-rank measure.
The second study scored 4,953 bioRxiv preprints from April 2025. QED then tracked where those papers were eventually published.
Researchers matched 2,879 preprints to published versions with an associated journal ranking. QED reported a Spearman correlation of 0.63 between its preprint scores and eventual journal rank.
The correlation varied across 21 disciplines. It reached 0.78 in genetics but fell to 0.39 in systems biology.
Those differences deserve attention. A single cross-disciplinary percentile can conceal meaningful variation in model performance, evidence conventions, and publication behavior.
The third study focused on disagreements. QED created 100 paper pairs where its score preferred the paper published in the lower-ranked journal.
Fifteen domain experts reviewed the pairs without seeing the QED Scores or publication venues. They produced 70 confident judgments, including 60 decisive choices after ties were excluded.
Experts selected the QED-favored paper in 75% of those decisive cases. The reported 95% confidence interval ranged from 63% to 84%.
That result supports QED’s argument that venue prestige can misrepresent paper-level quality. It does not establish that QED can identify truth without human judgment.
Publication venue is also an awkward benchmark. Journal placement depends on novelty, editorial scope, timing, presentation, reviewer availability, and strategic submission choices.
A score correlating with venue rank can validate alignment with existing publication outcomes. It cannot, by itself, validate a claim that the system escapes the publication system’s biases.
QED acknowledges an important boundary. Its methodology says The 1% is not a substitute for peer review.
That is the sensible framing. QED can prioritize manuscripts for examination, but a high percentile should not authorize clinical decisions or settle scientific disputes.
The primary contest is therefore direct assessment versus prestige-based inference. QED asks whether readers should inspect claims and evidence instead of borrowing a journal’s reputation.
That challenge is useful even if QED’s score remains imperfect. Journal branding has always compressed many judgments into a shorthand that readers can misuse.
A carefully validated AI score might become another signal. It should not become the only signal, or a new prestige label with less public accountability.
What the top 1% claim does not prove
QED’s internal results justify serious testing, but they do not yet justify unconditional trust.
The main limitation is independence. QED Science developed the system, designed the validation, published the methodology, and reported the headline performance figures.
Company-led studies can provide valuable evidence. Independent replication becomes essential when a product’s commercial value depends on the same performance claims.
One external review identified weaknesses across all three validation studies. The critique was produced through another AI-based research evaluation service, so it is not a definitive human replication.
Still, its questions are concrete. The review report argues that QED’s studies share related reference signals rather than offering fully independent confirmation.
The first study depends on expert-assigned quality categories. The second uses journal rank as an outcome. The third selects cases where QED and journal rank strongly disagree.
That structure can demonstrate consistency under QED’s chosen tests. It leaves open how the system performs in prospective decisions defined by outside researchers.
The second study also matched only 2,879 of 4,953 preprints to published versions with journal-ranking data. That leaves 2,074 manuscripts outside the reported correlation analysis.
Some unmatched papers might publish later, appear in venues without the required metric, or never reach formal publication. Their exclusion can change the apparent relationship.
Selection effects matter because publication is not random. Strong work can remain unpublished, while weak work can survive review.
The third study offers compelling disagreement data, but it uses an intentionally unusual sample. Researchers selected pairs because QED and journal rank pointed in opposite directions.
That design tests a relevant conflict. It does not estimate how often QED makes the better choice across ordinary papers.
The 75% preference result also comes from 60 decisive judgments. That is informative, but small compared with the 57,455 manuscripts featured in The 1%.
The selected list introduces another issue. “Top 1%” sounds objective, yet it remains relative to QED’s corpus, category handling, scoring version, and chosen evaluation dimensions.
Originality and validity are valuable criteria. They do not cover every quality that scientists care about.
A technically valid paper can be narrow or unimportant. An original paper can rely on fragile measurements. A replication can offer limited originality while providing enormous scientific value.
Ethics, data availability, methodological transparency, reproducibility, practical importance, and conflicts of interest can also affect how readers should use a result.
QED’s agents examine several related factors, but outsiders need more visibility into their weighting. They also need error analyses, field-level calibration, and examples of false positives.
Model opacity raises operational questions. QED says its architecture combines multiple large language models with proprietary models.
Its public methodology does not identify every underlying model, prompt, component weight, or update schedule. Those details can affect reproducibility.
If a model provider changes a system, the same manuscript might receive a different result. QED would need versioned scoring and stable audit records to support consequential use.
Research has already shown that LLM evaluators can reproduce social biases. A large economics experiment generated 27,090 evaluations across 9,030 papers.
That peer-review experiment found that models distinguished quality but favored prominent institutions, male authors, and renowned economists.
QED’s blind input should reduce several of those effects. It does not answer whether its model learned other preferences for familiar methods, popular topics, or conventional argument structures.
The system could also reward manuscripts that are easier for models to parse. Clear claim structure is good, but readability must not become an invisible substitute for evidentiary strength.
Adversarial behavior is another concern. Once authors understand what a scoring system rewards, some will optimize manuscripts for the metric.
That pattern appears wherever rankings distribute attention. Search engines, university metrics, citation measures, and journal impact factors have all encouraged strategic behavior.
A public score therefore needs resistance to manipulation. It also needs monitoring for hidden prompts, fabricated citations, duplicated evidence, and stylistic tactics designed to influence an evaluator.
Researchers should treat QED as a discovery aid until those questions receive independent answers. A high score can justify reading a paper sooner, not believing it sooner.
The reverse is equally important. A low score should not quietly bury unconventional research without an appeal path or a clear explanation.
For knowledge workers tracking scientific claims, maintaining source context matters more than collecting rankings. A searchable knowledge base can preserve papers, critiques, updates, and conflicting evidence together.
That workflow keeps the score in its proper role. It becomes one filter attached to a source, not a verdict detached from the underlying work.
Three signals will determine whether researchers should trust it
Trust will depend on independent validation, stable performance across fields, and evidence that scientists use the score without surrendering judgment.
The first signal is a prospective, preregistered study run by researchers outside QED Science. Independent teams should define the evaluation criteria before seeing the results.
They should test fresh manuscripts that neither the models nor evaluators have encountered. The study should include published papers, unpublished papers, replications, negative results, and intentionally flawed submissions.
External experts should assess claims and evidence without receiving QED’s labels. Researchers can then compare QED with human panels, journal decisions, replication outcomes, and later corrections.
No single benchmark provides perfect ground truth. Agreement across genuinely independent measures would strengthen QED’s claim more than another company-led comparison.
Failure under those conditions would weaken the assertion that QED offers a generally reliable measure of scientific quality. It might still remain useful for narrower triage tasks.
The second signal is transparent field-level performance over time. QED already reports correlations ranging from 0.39 to 0.78 across disciplines.
Future releases should disclose false-positive patterns, false-negative patterns, calibration drift, and scoring changes for each field. Aggregate accuracy can hide weak performance in specialized areas.
Versioning is especially important. Every score should identify the system version, reference corpus, evaluation date, and uncertainty around the ranking.
Researchers also need explanations that connect scores to specific claims and experiments. A percentile without traceable reasoning cannot support meaningful correction.
If QED publishes stable, reproducible field-level results, its case becomes stronger. If scores shift silently after model changes, researchers should limit the tool to informal discovery.
The third signal is how institutions use the ranking. Reading recommendations create lower stakes than funding screens, hiring filters, or clinical evidence decisions.
A tool that helps scientists discover overlooked papers can reduce prestige bias. The same tool can create algorithmic prestige if institutions treat the score as a cutoff.
Watch whether funders and journals use QED alongside expert review or before it. Also watch whether authors can challenge errors and obtain rescoring after revisions.
The healthiest adoption would preserve human responsibility. Researchers would use QED to identify papers, inspect its reasoning, and then evaluate the underlying evidence.
The riskiest adoption would hide decisions behind a percentile. An institution could reject work automatically while claiming the process was neutral because names were removed.
Nature’s question has a conditional answer. Researchers should trust QED enough to test its recommendations, but not enough to outsource scientific judgment.
QED has presented a credible response to a real bottleneck. Its scale, blind scoring, and claim-level analysis deserve careful independent examination.
The Horizon Nature debate now needs evidence from outside the company. Over the next few months, look for preregistered replication, field-specific error reporting, and disclosed institutional use.
Would a QED score change which preprint you read first? It probably should. Would it change what you believe without checking the methods and evidence? It should not.



