Skip to content

Benchmarks

A public benchmark of AI evidence retrieval and fidelity.

There is currently no shared, public way to answer a basic question: when a system answers a scientific question, does it reach the right evidence, and does it describe that evidence accurately? This page describes a benchmark intended to fill that gap, and states plainly that it does not exist yet.

Current status

The benchmark does not exist yet

No question set has been published, no results have been produced, and no system has been evaluated. Nothing on this page should be read as a measurement, a comparison, or a claim about any product. It is a specification of intent, written down early so it can be held to.

Scope

Six things it would measure separately.

Most of the value is in keeping these apart. A single composite number would obscure exactly the failures that matter.

01Retrieval
Given a question with a known adequate answer in the literature, is the adequate source reached at all? Reported as a rate with an error estimate, per field and per age band.
02Citation
Of the sources reached, which are actually named in the answer? Retrieval and citation diverge constantly, and a benchmark that conflates them measures neither.
03Grounding
Does the cited source support the specific claim it is attached to? A correct citation attached to the wrong sentence is a distinct and common failure.
04Fidelity
Does the answer preserve the study's population, endpoint, effect size, and stated limitations, or does it quietly broaden them? Overstatement is the failure mode with the most direct clinical consequence.
05Recency and correction handling
Is superseded, corrected, or retracted evidence still being used? Whether a system respects a retraction is one of the sharpest tests available.
06Stability
Asked the same question five times, does the same evidence appear? An unstable answer is not a ranking, and reporting a single run as if it were one would be misleading.

Credibility

What would make it worth believing.

A benchmark from a commercial party is worth exactly as much as its verifiability. These are the conditions under which it should be taken seriously, and by the same token the conditions under which it should be dismissed if they are not met.

01A public, versioned question set
Questions, reference answers, and adequacy judgements published in full, with a version identifier, so any figure can be re-derived by someone who does not trust the party that produced it.
02Documented adjudication and agreement
Human judgement decides what counts as an adequate source and a faithful summary. The rubric, the annotator qualifications, and inter-annotator agreement must be published, or the numbers mean nothing.
03Honest labelling of what was measured
An API model with web search is not a vendor's consumer product. Any result must state exactly which surface, which configuration, and which date window produced it, and must not be presented as a consumer product ranking.
04Independent replication
A benchmark published by a company that sells visibility measurement has an obvious conflict of interest. It is credible only if an independent group can and does reproduce it, which requires releasing enough for them to try.
05Error bars and sampling design
Sample sizes, sampling frame, and confidence intervals reported with every figure. A leaderboard of point estimates without variance is a marketing artefact.
06External methodology review
Changes to the instrument reviewed by researchers with no commercial interest in any particular result, and a published changelog when the instrument changes.
07Published negative results
The measurements that did not work, and the cases where the benchmark failed to distinguish systems, published alongside the ones that did.
08Resistance to optimisation
A held-out portion of the question set, rotated over time, so the benchmark measures capability rather than familiarity with the benchmark.

Boundaries

What it would deliberately not be.

  • A leaderboard of AI products.
  • A quality score for journals, articles, or researchers.
  • A claim about any vendor's consumer product based on an API measurement.
  • A commercial input to the product's own visibility scores.

Conflict of interest, stated up front

Trace sells evidence-visibility measurement. A benchmark it publishes about evidence retrieval is therefore not disinterested. That is the reason for the independence and replication conditions above, and it is the reason none of this will be presented as settled until someone unaffiliated has reproduced it.