Measures single-run greedy-decoding performance for each model before and after training.
Proprietary models were evaluated with the same code and the same grading, using 8,192 tokens per turn. The two max-effort runs used 32,768 tokens to avoid truncating longer reasoning.