Skip to content

All resources

Guide6 min

What makes a title and abstract machine-readable

Concrete, unglamorous properties that make a record easier to match — none of which require writing worse science.

Full text

'Machine-readable' is a poor phrase for what actually matters here. Nothing is being read in the sense of parsing a schema. What is happening is closer to matching: a question is compared against your text, and the more clearly your text states what the work was about, the more reliably that comparison succeeds. The properties below are the ones that do most of the work, and almost all of them are things a careful editor would ask for anyway.

Titles

State the finding or the question, not the theme. 'A randomised trial of X in adults with Y' tells a matcher the design, the intervention, and the population. 'Revisiting X: new perspectives' tells it almost nothing, and tells a human reader little more.

Include the population and the intervention explicitly. Terms that appear only in the body are far weaker signals than terms in the title. If your study is in older adults with moderate disease, the title is the cheapest place to say so.

Avoid decorative constructions. Colons carrying a rhetorical flourish, questions with no stated subject, and puns all reduce the density of matchable content in the highest-weighted field you have. This is not an argument for dullness; it is an argument for spending the title's limited space on the study rather than on the framing.

Prefer the term your readers would use, and include the technical synonym where both exist. If a condition is known by a common name and a formal one, a title carrying both is not redundant — it is the only place the two get connected for a system that has no other way to know they are the same thing.

Abstracts

Structure is the single largest lever. Population, intervention, comparator, outcome, and principal limitation should each be identifiable, whether through explicit headings or through a paragraph structure that keeps them separable. Structured abstracts became standard in clinical publishing because they help human readers find the same five things quickly; the machine benefit is a side effect of the same clarity.

State the population precisely, in the first two sentences. Sample size, setting, age range, and disease severity are the details that determine whether your paper is the right answer to a specific question. Buried in the middle of a narrative paragraph, they are much less likely to be matched against.

Name the outcome measure explicitly, with the instrument if there is one. 'Improved function' is unmatched; 'change in six-minute walk distance at twelve weeks' is a concrete handle that also happens to be honest about what was measured.

Give the direction and the size of the effect, with its uncertainty. An abstract that reports a result without a magnitude forces any downstream summary to characterise it in vague terms, which is precisely how overstatement enters the record. If your abstract says how large the effect was and how uncertain, a summary that widens it is visibly wrong rather than merely unsupported.

State the principal limitation in the abstract, not only in the discussion. Summaries are frequently built from the abstract alone. A limitation that appears only on page eleven will not survive into any of them, and its absence will look like a claim you did not make.

Spell out every abbreviation on first use, including the ones that feel universal within your field. An acronym that means one thing in cardiology and another in software is a matching hazard.

The record around the text

Register the identifiers and make sure they resolve. An unregistered or malformed identifier breaks the links that let anything connect your work to its citations, its corrections, and its versions.

Link the preprint to the version of record, in both directions. Unlinked versions split a single work into several weakly-connected records, each carrying a fraction of the signal.

Use controlled vocabulary terms if your venue supports them, and choose them for the questions your work answers rather than for topical breadth. Three precise subject terms are worth more than ten aspirational ones.

Attach trial or protocol registration identifiers where they exist, and make sure the funding acknowledgement is in the structured field rather than only in the text.

The line not to cross

Everything above is a request to be more specific. None of it is a request to repeat terms, to inflate claims, to add populations you did not study, or to write a title engineered for matching rather than accuracy. Those tactics degrade the record, mislead readers, and tend to be short-lived against systems that improve. The version of this work that lasts is the version that makes the paper easier to understand correctly.

More guides