Reach vs. Recall@100
share of gold references (%) · +N = GPT-5.5 lead in points (− = GPT-5.1 lead)
must-cite · overall
macro mean (n=50, gold = 884 refs)
must-cite overall. recall@reach 5.1 60.0, 5.5 63.9. recall@100 5.1 42.8, 5.5 57.0 (percent).
related-work · overall
macro mean (n=50, gold = 969 refs)
related-work overall. recall@reach 5.1 52.3, 5.5 56.4. recall@100 5.1 37.3, 5.5 47.9 (percent).
must-cite · reach
reached during 20 iterations
must-cite reach. 5.1: 36.4, 44.8, 51.8, 56.6, 68.7, 69.9. 5.5: 40.6, 49.3, 55.3, 64.5, 71.8, 74.2 (percent).
paper publication year
must-cite · recall@100
kept in the top 100
must-cite recall@100. 5.1: 29.4, 35.8, 42.4, 39.5, 49.3, 44.5. 5.5: 39.2, 47.8, 49.4, 57.2, 62.6, 64.1 (percent).
paper publication year
GPT-5.1 GPT-5.5
Figure 7. Reach versus recall@100 for GPT-5.1 and GPT-5.5 on the literature-review task.