OASYS @ MIT

identical_pairs_repr — closest REPRESENTATIVE eval↔train RLM trajectory per case

Exactly 9 samples — one closest eval↔train trajectory pair per RLM case (6 length-gen tasks + 3 strategy-gen panels). Each is representative: both sides reward>0.5, trained checkpoints, and eval checkpoint ≥ train checkpoint (the model could only have trained on that sample before producing the eval).

Method

Results (token_lcs = behaviour-view selection score; all satisfy eval_step ≥ train_step)

case verdict lined-up token_lcs token_lev eval@step (r/turns) train@step (r/turns) dir
OOLONG NEARLY_IDENTICAL 93/100 0.824 0.723 @131 (r1.00/8) @118 (r1.00/8) OOLONG/
LongBenchPro NEARLY_IDENTICAL 95/100 0.767 0.665 @131 (r1.00/4) @43 (r1.00/4) LongBenchPro/
GraphWalks NEARLY_IDENTICAL 93/100 0.621 0.476 @121 (r1.00/6) @118 (r1.00/6) GraphWalks/
OOLONG-Pairs MOSTLY_SAME 84/100 0.628 0.469 @91 (r0.91/6) @54 (r0.81/6) OOLONG-Pairs/
AdaLEval-BestAnswer SAME_STRATEGY 84/100 0.564 0.413 @111 (r1.00/6) @87 (r1.00/7) AdaLEval-BestAnswer/
MRCR SAME_STRATEGY 82/100 0.667 0.518 @61 (r1.00/7) @43 (r1.00/5) MRCR/
SG-OBLIQ-twitter-wildchat SAME_STRATEGY 86/100 0.599 0.470 @181 (r1.00/6) @172 (r0.63/6) SG-OBLIQ-twitter-wildchat/
SG-OOLONG-trec-spam NEARLY_IDENTICAL 92/100 0.514 0.360 @161 (r1.00/5) @150 (r1.00/5) SG-OOLONG-trec-spam/
SG-OBLIQ-writing-math NEARLY_IDENTICAL 93/100 0.494 0.314 @281 (r1.00/5) @266 (r0.57/5) SG-OBLIQ-writing-math/

Per-case pair identities

Notes

Similarity figures (../../plots/) — causal (train at/before eval), TF-IDF dropped