The benchmark is contaminated, and it must be said out loud
The median causal gene already carries every one of its patient's terms, because HPO's gene annotations are curated from these same case reports. This measures how well the ranking retrieves a gene HPO has already been told about: an upper bound. A prospective number, on a patient whose gene nobody has annotated yet, would be lower, and this corpus cannot say by how much.
Measuring my own changes, including the one that failed
Information-content weighting and ontology propagation both replaced plain term counting. Asked whether they helped, the corpus said only one of them did — pessimistic figures, 6,485 cases with six or more terms:
| count terms (original) | 61.8% | 0.682 |
| + information content | 63.6% | 0.706 |
| + propagation | 58.2% | 0.653 |
| + both (shipped) | 59.5% | 0.670 |
Weighting earns its place. Propagation costs about what weighting gains — kept for a reason the documentation argues rather than assumes, with the table there so a reader can disagree.