Why retrieval, citation, and evidence use are different things
Three measurements that get collapsed into one number, and what is lost each time it happens.
Full text
When people ask whether a paper is 'visible' in AI-mediated discovery, they are usually asking three questions at once and hoping for a single answer. The three questions have different mechanisms, different failure modes, and different remedies. A number that averages them tells you that something is wrong without telling you what, which is the least useful diagnostic shape available.
Retrieval: was it reached?
Retrieval asks whether the paper entered the candidate set at all. A system given a question gathers some documents; retrieval measures whether yours was among them.
Failure here is almost always about matching. The vocabulary of the question does not meet the vocabulary of the record, the abstract offers little structured surface, the identifiers are incomplete, or no open version exists to be indexed. These are metadata and phrasing problems, and they are the most fixable class of failure in the entire chain.
Retrieval is also the only one of the three that is a strict prerequisite. If a paper is never retrieved, nothing downstream can happen. That makes it the first thing to measure and the first thing to fix.
Citation: was it named?
Citation asks whether, having been retrieved, the paper is actually referenced in the answer. Retrieved-but-not-cited is extremely common, and it is a genuinely different situation from not being retrieved.
A paper can be reached and then passed over because another source states the same finding more directly, because it is older than something else in the candidate set, because its abstract does not make its contribution legible in the first few lines, or because the system had room for three citations and yours was fourth. Some of these are addressable by clearer framing; some are competitive facts about the literature that no amount of metadata work will change.
The diagnostic value lies in the contrast. A paper with high retrieval and low citation has no discoverability problem — it has a positioning problem, and the remedy is about how clearly the contribution is stated, not about keywords or deposits. Averaging the two would hide exactly that distinction.
Evidence use: did it actually support the claim?
The third question is the one most often skipped, because it is the hardest to measure. A citation can be attached to an answer without the cited work supporting the specific sentence it sits next to. Conversely, an answer can rest substantially on a paper's finding while citing something else entirely — a review, a later paper, a secondary source that repeated the result.
Evidence use asks what the answer's claims actually rest on. It requires comparing the assertion against the source's own stated population, endpoint, and effect, and judging whether the source bears the weight. This is annotation work, it needs a published rubric and inter-annotator agreement statistics to mean anything, and it cannot be automated into a clean scalar without losing the thing that makes it valuable.
It is also where the most consequential failures live. A paper that is retrieved, cited, and then described as supporting a broader population or a firmer effect than it studied has been used to make the record worse. That is a fidelity failure, and it is invisible to any measure that stops at citation counts.
Why the composite is still tempting
A single number sorts. It fits in a column, it goes on a dashboard, and it makes a portfolio comparable at a glance. There is real operational value in that, which is why composites keep getting built.
The failure mode is well documented from the history of bibliometrics: a composite built for internal triage escapes into evaluation, where it is used to make decisions about people and journals that its construction never supported. The defence is procedural rather than clever. Report the components separately and prominently. Attach the version of the method to every number. State that the composite is a heuristic under validation. And decline to publish it as a public quality mark, because once a number is displayed as a badge, its stated caveats stop travelling with it.
In practice, the components are also just more useful. 'Not retrieved for the questions it answers' and 'retrieved constantly but never the source that gets cited' are two different problems with two different next actions. A composite tells you neither.
More guides
What AI-mediated discovery changes for authors
The path from a published paper to a reader has moved. What that means in practice for the person who wrote it.
What makes a title and abstract machine-readable
Concrete, unglamorous properties that make a record easier to match — none of which require writing worse science.
How to read a synthetic stakeholder panel responsibly
What a simulated panel can legitimately tell you, what it cannot, and the specific ways it goes wrong.