Biomedical evidence evaluation

Read the answer with its evidence.

Independent original analysis of PubMedQA and BioASQ 12b: datasets, scoring, evidence phases, version caveats and reproducible biomedical QA comparisons.

Biomedical QA / evidence availability

What is supplied, and what must be found?

Supplied evidence

PubMedQA

Question + abstractYes / No / Maybe

500 questions in held-out evaluation [1]

The main task withholds the abstract conclusion.

Evidence changes by phase

BioASQ 12b

  1. A

    retrieve

    Return literature documents and snippets.

  2. A+

    answer with retrieved evidence

    Generate answers before receiving gold evidence.

  3. B

    answer with gold evidence

    Answer with manually selected supporting material.

340 test questions · Yes/no, factoid, list and summary. [5][6][7]

01 /

Input boundary

PubMedQA’s withheld conclusion changes the task definition.

02 /

Evidence condition

BioASQ A+ and B distinguish system-found evidence from gold evidence.

03 /

Run identity

Keep edition, population, parser and grader attached to every result.

Our analytical question

We analyze two named biomedical question-answering benchmarks at the level of their inputs, references and scoring. PubMedQA isolates answering from a supplied abstract; BioASQ 12b separates finding literature from answering with it. Explore the published task conditions, inspect selected historical measurements and use our original guides to make an evaluation reproducible. Benchmark authors retain credit for the datasets and experiments. Our contribution is the analysis and tools; we have not run the models displayed here.

Read the original evidence closely

The benchmark, unpacked.

Primary source library ↗
Dossier01

Original EMNLP 2019 reasoning-required protocol

PubMedQA ↗

One abstract. Three labels. A surprisingly important boundary.

UnitOne research question / abstract pairMeasureAccuracy and macro-F1
Dossier02

2024 challenge, frozen edition

BioASQ 12b ↗

Separate finding evidence from answering with evidence.

UnitOne biomedical question in one challenge batch and phaseMeasureTask-specific metrics

An original analytical tool

Explore the biomedical evidence contract

Evidence explorer

Filter actual benchmark conditions by how evidence is supplied and how answers are judged. The annotations are our original comparison, not new performance measurements.

8 of 8 evidence entries shown

Supplied evidence

PubMedQA: question + abstract

Read dossier ↗
Input
Non-conclusion abstract and research question
Output
Yes/no/maybe
Measure
Accuracy; macro-F1
Interpretation boundary

Conclusion must remain withheld at test time.

[1][3]
Supplied evidence

PubMedQA: question + conclusion

Read dossier ↗
Input
Question and the abstract conclusion
Output
Yes/no/maybe
Measure
Accuracy; macro-F1
Interpretation boundary

Reasoning-free condition, not directly interchangeable with the main benchmark.

[1]
Retrieved evidence

BioASQ: document retrieval

Read dossier ↗
Input
Biomedical question
Output
Ranked relevant articles
Measure
MAP
Interpretation boundary

Retrieval reference enrichment and batch must match.

[7][5]
Retrieved evidence

BioASQ: snippet retrieval

Read dossier ↗
Input
Biomedical question
Output
Relevant text snippets
Measure
F-measure
Interpretation boundary

A document hit and a snippet hit have different units.

[7]
Exact answers

BioASQ: yes/no answers

Read dossier ↗
Input
Question plus phase-permitted evidence
Output
Yes or no
Measure
Accuracy
Interpretation boundary

A+ and B differ in evidence availability.

[7][5]
Exact answers

BioASQ: factoid answers

Read dossier ↗
Input
Question plus phase-permitted evidence
Output
Ranked exact answers
Measure
MRR
Interpretation boundary

Ranking an answer is distinct from producing a summary.

[7]
Exact answers

BioASQ: list answers

Read dossier ↗
Input
Question plus phase-permitted evidence
Output
Set of answer entities
Measure
Mean F-measure
Interpretation boundary

Coverage and extra entries both matter.

[7]
Prose answers

BioASQ: ideal answers

Read dossier ↗
Input
Question plus phase-permitted evidence
Output
Explanatory prose answer
Measure
Mean manual score
Interpretation boundary

Manual quality judgment is not exact-answer accuracy.

[7]

This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3][4][5][6][7]

Original analysis / methods and interpretation

What the score leaves unsaid.

All analyses →

Questions, answered

Read the result in context.

Specific tasks. Stated conditions.
Inspect every source.

What does PubMedQA measure?

It measures yes/no/maybe answers to research questions using corresponding abstracts. Its main reasoning-required setting withholds the abstract conclusion at test time.

How is BioASQ 12b different?

The 2024 challenge separates literature retrieval, answering from system-retrieved evidence and answering with gold evidence. It includes several answer types with different scoring rules.

Are these original benchmarks created by Med Evals?

No. They are established benchmarks credited to their creators. Our original work is the analytical interpretation, coverage explorer and implementation guidance.

Why are two training totals shown for BioASQ 12b?

The official catalog and overview prose report 5,046 questions, while the paper’s Table 1 reports 5,049. We preserve that discrepancy and recommend a versioned manifest for a reproducible run.

Working tool / saved on this device

Prepare a benchmark comparison brief

Interactive worksheet

Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.

Fix the evidence contract

Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.

Download the evidence ↗