Model Performance After Training

Arcwise-Plat-SQL · 498 questions · Trained on River
Base model (8,192 tokens/turn) After RL
Gain Vs human Avg tokens / question
base → RL
Avg token
change
Arcwise accuracy (%)

Measures single-run greedy-decoding performance for each model before and after training.

Proprietary models were evaluated with the same code and the same grading, using 8,192 tokens per turn. The two max-effort runs used 32,768 tokens to avoid truncating longer reasoning.