Skip to content

Discover

Know whether the right evidence is being found.

A paper can be highly cited in the literature and absent from the questions people now ask machines. Trace Discover measures the path from question to answer: which questions should reach your work, which engines retrieve it, which answers cite it, and whether the finding arrives intact.

Query universe

Start from the questions the work should be able to answer.

Visibility is meaningless without a defined denominator. Every measurement begins with a versioned set of questions derived from the paper's population, intervention, comparator, outcomes, and stated limitations — then reviewed and edited by you before any run.

Clinical decision questions
Example — What is the evidence for velunemab in adults with Novera syndrome?
Mechanism questions
Example — How does velunemab affect the pathway implicated in Novera syndrome?
Comparative questions
Example — How does velunemab compare with standard care on the primary endpoint?
Safety questions
Example — What adverse events were reported in trials of velunemab?
Population questions
Example — Was velunemab studied in patients over 65 with renal impairment?
Evidence-strength questions
Example — How strong is the evidence base for velunemab, and what are its limits?

Examples use the fictional TRACE-101 trial from the demo workspace. Question sets are versioned, so a change in score can be attributed to a change in the world rather than a change in the instrument.

Multi-surface measurement

One paper. Several very different readers.

Engines differ in what they index, how they retrieve, and how much of a source they read. A paper strong on one surface can be invisible on another, so results are reported per surface and never averaged into a single opaque figure.

Surfaces

Measured independently

  • General assistants

    Consumer-facing chat surfaces where a patient or journalist is likely to ask.

  • Search-grounded answers

    Answer engines that retrieve live sources and summarise them inline.

  • Research assistants

    Literature-specialised tools that read abstracts and full text directly.

  • Retrieval APIs

    Programmatic retrieval, sampled to separate the index from the model.

Runs record the engine, the date, and the methodology version. Model versions change without notice; an undated visibility number is not a measurement.

Retrieval vs citation vs use

Three different failures, three different fixes.

Being retrieved is not being cited. Being cited is not having your evidence used. Having your evidence used is not being understood. Each stage is detected separately, so a drop points at a cause.

Retrieved

62%

The paper entered the engine's retrieved source set for the question. It may never reach the reader.

Cited

43%

The paper is visibly referenced in the answer the reader actually sees, with an identifier or link.

Evidence used

31%

A claim in the answer is attributable to a finding in this paper, whether or not the citation is shown.

Interpreted faithfully

27%

Population, effect size, and stated limitations survived the summary without distortion.

Illustrative figures from the demo workspace, measured against the same question set. The gap between 62% retrieved and 43% cited is a surfacing problem; the gap between 31% used and 27% faithful is an interpretation problem.

Competitor papers

See which papers are answering your questions instead.

For every question in the set, the sources that were actually cited are recorded. Over a run this produces a ranked list of the papers occupying the evidence space around your work — often reviews, guidelines, or older trials rather than direct rivals.

Demo workspace

Frequently cited alongside TRACE-101

DEMO DATA
  • 10.5555/novera.review.2024

    Narrative review, cited on 41% of questions

  • 10.5555/novera.guideline.2023

    Society guideline, cited on 33% of questions

  • 10.5555/velunemab.phase2.2022

    Earlier phase 2, cited on 19% of questions

Fictional records created for demonstration. They are not real publications and must not be cited.

Longitudinal monitoring

A baseline is a starting point, not a verdict.

Re-runs on a fixed schedule keep the question set constant while the engines move underneath it. Results are stored as history with the methodology version attached, so old numbers stay interpretable.

Changes are described, not explained

A shift in retrieval after a model update is recorded as having occurred after that update. Trace does not infer causation from a coincidence in timing, and never promises that any action will improve a ranking.

Monitor in detail

How every number here is produced

Query generation, sampling, engine definitions, retrieval and citation detection, fidelity scoring, and the known limits of each are documented in full. If a number on this page cannot be traced back to a run, treat it as marketing rather than measurement.

Read the methodology

Measure one paper. Then decide what to change.