Original EMNLP 2019 reasoning-required protocol
PubMedQA ↗
One abstract. Three labels. A surprisingly important boundary.
Biomedical evidence evaluation
Independent original analysis of PubMedQA and BioASQ 12b: datasets, scoring, evidence phases, version caveats and reproducible biomedical QA comparisons.
500 questions in held-out evaluation [1]
The main task withholds the abstract conclusion.
Return literature documents and snippets.
Generate answers before receiving gold evidence.
Answer with manually selected supporting material.
340 test questions · Yes/no, factoid, list and summary. [5][6][7]
PubMedQA’s withheld conclusion changes the task definition.
BioASQ A+ and B distinguish system-found evidence from gold evidence.
Keep edition, population, parser and grader attached to every result.
We analyze two named biomedical question-answering benchmarks at the level of their inputs, references and scoring. PubMedQA isolates answering from a supplied abstract; BioASQ 12b separates finding literature from answering with it. Explore the published task conditions, inspect selected historical measurements and use our original guides to make an evaluation reproducible. Benchmark authors retain credit for the datasets and experiments. Our contribution is the analysis and tools; we have not run the models displayed here.
Read the original evidence closely
Original EMNLP 2019 reasoning-required protocol
One abstract. Three labels. A surprisingly important boundary.
2024 challenge, frozen edition
Separate finding evidence from answering with evidence.
An original analytical tool
Filter actual benchmark conditions by how evidence is supplied and how answers are judged. The annotations are our original comparison, not new performance measurements.
8 of 8 evidence entries shown
Reasoning-free condition, not directly interchangeable with the main benchmark.
[1]This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3][4][5][6][7]
Original analysis / methods and interpretation
An original engineering analysis of conclusion leakage, class balance and the official PubMedQA evaluator.
Use the A, A+ and B conditions to distinguish retrieval failures from answer-generation failures.
Compare abstract interpretation with literature retrieval and answer synthesis without inventing a combined leaderboard.
Questions, answered
Specific tasks. Stated conditions.
Inspect every source.
It measures yes/no/maybe answers to research questions using corresponding abstracts. Its main reasoning-required setting withholds the abstract conclusion at test time.
The 2024 challenge separates literature retrieval, answering from system-retrieved evidence and answering with gold evidence. It includes several answer types with different scoring rules.
No. They are established benchmarks credited to their creators. Our original work is the analytical interpretation, coverage explorer and implementation guidance.
The official catalog and overview prose report 5,046 questions, while the paper’s Table 1 reports 5,049. We preserve that discrepancy and recommend a versioned manifest for a reproducible run.
Working tool / saved on this device
Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.
Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.
Download the evidence ↗