Achieving Frontier Text-to-SQL with the River API
We post-trained open models that outperform proprietary frontier models on Arcwise-Plat, a text-to-SQL benchmark. Our results beat GPT-6 Astra Pro and Opus 5.5 for less than 1% of the cost.
We started from ReViSQL (opens in a new tab), which trains Kimi K2.6 with RL and a LoRA adapter. Our team rebuilt the recipe on River, fixed several implementation issues, and improved performance and training stability. You can find it in our open source recipe book (opens in a new tab).
With our recipe we were able to push four open models to human-level performance (92.96%) on Arcwise-Plat.
Results
Our trained models get the best of both worlds: frontier performance at a fraction of the cost. Five outperformed the strongest proprietary model we measured, GPT-6 Astra Pro at max effort, averaging 4.0 points higher at less than 2% of the cost per question.GLM-5.3 Flash led the group, scoring 4.6 points higher at just 0.5% of the cost.
Token Use After Training
Cost per question depends on both a model’s token price and how many tokens it uses. The larger MoE models became more concise, using fewer turns and less reasoning, which lowered their cost. The smaller models generally moved the other way, spending more time reasoning and exploring the database.
We break out input and output below because they’re priced differently and respond differently to training. Each new turn re-sends the conversation so far, so more SQL exploration can quickly drive up input tokens. Qwen3.6 shows this clearly, with its increase in database exploration leading to a large jump in input tokens.
Output tokens reflect how much the model reasons and are typically more expensive. Qwen3.8 saw only a modest increase in input, but its reasoning output more than doubled, giving it the largest cost increase among the models we measured.
Token Use per Question
Arcwise-Plat, base model → trained
| Model | Turns | Input tokens | Output tokens | Cost change |
|---|---|---|---|---|
| Kimi K2.6 | 3.4 → 3.1 | 11.4k → 10.6k | 2.4k → 1.4k | −25% |
| GLM-5.3 Flash | 3.3 → 3.0 | 11.5k → 10.3k | 2.0k → 1.3k | −19% |
| DeepSeek V4.1 Flash | 3.4 → 3.1 | 12.5k → 11.3k | 1.7k → 1.1k | −19% |
| Nemotron 3.5 Lightning | 2.7 → 3.7 | 9.6k → 13.7k | 1.8k → 2.6k | +42% |
| Qwen3.6-35B-A3B | 1.2 → 4.7 | 4.2k → 17.9k | 2.7k → 3.3k | +80% |
| Qwen3.8-27B | 2.5 → 3.2 | 8.9k → 13.2k | 0.9k → 2.1k | +131% |
Base models are measured with 8,192 tokens per turn and trained models with their recipe’s training budget. Costs use OpenRouter list prices as of Sep 29, 2026.
A few results stand out:
Greatest Performance Gain
Nemotron 3.5 Lightning
Training improved accuracy by 37.8 points with only a modest increase in cost. It remained the most cost-efficient model we tested.
Most exploration
Qwen3.6-35B-A3B
Training shifted it from the least exploratory model to the most, from 1.2 to 4.7 turns per question.
Largest cost increase
Qwen3.8-27B
Driven primarily by a sharp rise in output tokens.
Cost ranking
DeepSeek V4.1 Flash
Became cheaper than both Qwen models, while the rest maintained their relative order by OpenRouter cost.
Results Across Benchmarks
Because we tuned the recipe and selected checkpoints on Arcwise-Plat, we also wanted to know whether the gains transferred beyond that benchmark. We tested the same models on Spider2-Lite, a separate benchmark with 135 harder SQLite tasks, using Spider2’s own comparator.
Every model improved after training, suggesting the gains carried beyond Arcwise-Plat rather than coming at the expense of performance elsewhere. Spider2-Lite proved a tougher benchmark, with our best models finishing just shy of the proprietary frontier.
Spider2-Lite Results
How We Evaluate
We evaluate Arcwise-Plat every five training steps. Each checkpoint is tested on all 498 questions with one greedy attempt per question and graded the same way as ReViSQL.
For each model, we report the best checkpoint evaluation over the run.
Our best Qwen3.6-35B-A3B run
The gold ring marks the run's best Arcwise evaluation, the result we report; the teal ring marks the checkpoint with the best validation score.
The ReViSQL authors use peak validation accuracy to select a checkpoint for evaluation, but validation quickly lost resolution in our experiments. In our best Qwen3.6-35B-A3B run, validation had largely leveled off by step 25 and remained essentially flat for the rest of training. Arcwise-Plat performance, meanwhile, continued to improve along a logarithmic curve, gaining as much as 5.2 points from step 25 onward.
Given validation’s saturation, we had to turn to a different strategy to select our strongest checkpoint. We ran Arcwise-Plat at regular intervals throughout training and selected the highest-scoring checkpoint from those intervals. We chose every 5 steps as a good trade off of granularity vs sampling efficiency.
We also checked that these peaks were not isolated artifacts. Every champion checkpoint was within one point of peak validation accuracy, and the strongest checkpoints consistently outperformed their neighbors when re-evaluated. We explore these results further in the reflection on results section below.
A note on grading. Every result we report is graded under the ReViSQL authors’ rules: the same corrected reference queries they grade against, and their 1% tolerance on single-number answers. We use these rules for reported results so our numbers can be compared directly with theirs. Training and validation use our own grader, which is stricter and requires numeric answers to match exactly. Under our own grader, our best Arcwise-Plat scores average about half a point lower, and the other models we evaluated would shift to some degree as well. Matching the authors’ grader simplified our comparisons.
Next, we break down the recipe so you can better understand how we got these results. We also explore which parts of it had the most impact and the extensions we experimented with.
The ReViSQL Recipe
ReViSQL trains with reinforcement learning, repeating the same basic loop:
Sample
Give the model a batch of text-to-SQL questions and have it attempt each question 16 times. On each attempt, the model reasons through the problem and can query the database up to five times. It uses those results to produce a final SQL query.
Question BatchScore
Execute the final query and score it against the gold result. ReViSQL uses VeriEQL (opens in a new tab) to catch queries that are only accidentally correct, and also penalizes attempts that ignore evidence provided with the question.
AnswersCompare
For each question, compare the model's attempts against one another. Better-than-average attempts are reinforced, while worse ones are discouraged. If every attempt gets the same score, there’s nothing to learn from the comparison, so the question provides no training signal.
Group average: (2 − 1) / 8 = 0.125✓ Correct: 1 − 0.125+0.875× Incorrect: 0 − 0.125−0.125– No solution: −1 − 0.125−1.125Update
Train on the model’s entire attempt, including its reasoning, exploratory queries, and final answer. The database responses provide context, but aren’t trained on. Over time, successful approaches become more likely and unsuccessful ones less likely.
selectSuccessful attempts raise token probabilities. Unsuccessful attempts lower them.
Each cycle follows a different question. Real updates depend on the full attempts and training objective; shaping penalties and tokenizer details are omitted.
The idea is simple: generate several candidate solutions to the same problem, test which ones actually work, and train the model to favor those approaches.
The Authors’ Code
The ReViSQL authors had released their training code (opens in a new tab), so we used it as our reference implementation on River. In our initial tests, longer runs consistently collapsed around the second or third epoch, making the reported result difficult to reproduce.
Since the training code was originally written for Tinker, we first wanted to see whether differences between training APIs affected the result. We tested this by running the same author implementation across River, Tinker, and Fireworks.
Unmodified ReViSQL trainer example and data, 33 steps (1 epoch). Greedy evaluation every five steps.
The authors’ example defaults to one epoch, or 33 steps, which we used for this comparison. Every run fell short of the paper’s reported Kimi K2.6 result. That wasn’t too surprising, since the paper itself recommends training beyond one epoch. River’s stronger first-epoch result gave us confidence that the platform difference was not the source of the gap.
The method itself still looked sound, but the released code was not giving us a reliable path to the reported result. So we got to work and rebuilt the training loop from the ground up around the method described in the paper.
Our Implementation
Rebuilding the training code from scratch gave us a cleaner foundation for longer experiments and made it easier to isolate what was actually affecting performance. The following are some highlighted changes we made in our code to improve performance and stability.
Async RL
The released loop waits for every rollout in a batch before training. Because text-to-SQL rollouts vary widely in length, slow groups can leave training capacity idle.
Our implementation overlaps sampling and training, keeps question groups intact for reward comparison, and limits how stale a rollout can be. This improved throughput while preserving comparable accuracy at matched steps.
In practice, our recipe's recommended settings take about half the time of the original synchronous loop on the same API. This faster feedback loop allowed us to test far more variants of the recipe.
Learn more about how you can apply async RL in your own runs in our docs (opens in a new tab).
Correctness and reproducibility
Grader
We sharpened the grading check inside the training loop. The authors’ grader allowed a 1% tolerance when a query returned a single numeric value. This simplifies floating-point comparisons, but also lets incorrect row-count queries receive full credit when their results are close enough to the gold answer. We tightened this check so small counting errors would not be rewarded as correct.
Kimi Response Parsing
When Kimi K2.6 hit the token limit mid-reasoning, the authors’ released code treated the unfinished reasoning as its answer and graded it. That reasoning sometimes held a draft query, so an attempt that ran out of tokens could still be graded correct.
This sent the wrong training signal. An unfinished attempt could earn the same reward as a completed one, so the model never learns to wrap up reasoning. Worse, long attempts could still be reinforced whenever a draft query happened to be correct. This would slowly lead to runaway degenerate generations on most runs.
In our implementation, an interrupted attempt has no answer and receives a −1 reward, the lowest possible score. That makes finishing within the budget strictly better.
We suspect this bug was one of the main reasons we couldn’t reproduce the reported Kimi results with the authors’ code.
Outdated VeriEQL
The authors’ released code pinned a version of VeriEQL from before a fix for a known bug with ORDER BY inside subqueries, so it “refuted” some correct queries and cost them reward. That penalty fell on a pattern used in about 3% of the training questions’ reference queries, such as picking the top row inside a subquery. We use the fixed version.
Tool-Use Demonstration
Our implementation prepends a short worked example of a tool call and final answer to the conversation history. Without it, models still follow the required format, but early responses are much longer. For example, Kimi K2.6 used about 1,000 extra tokens per answer and hit the token limit far more often at the start of training. Ablating the example did not meaningfully affect accuracy, so we keep it because it makes early training more efficient.
Optimizing for longer runs
Validation accuracy often peaked earlier than Arcwise accuracy (in five of our six strongest runs), so the late steps were chasing small improvements, and a constant learning rate proved too aggressive for them. We switched to a schedule that warms the learning rate up over the first steps and then decays it along a cosine curve, locking in early gains while holding the peak for longer.
Now that we had a reliable implementation of the recipe, we got to work probing the claims in the paper and refining the recipe for each engine on our API.
Recipe Insights
Many parts of the ReViSQL recipe had a clear impact on performance. A few stood out across our experiments.
The BIRD-Platinum Dataset
BIRD-Platinum (opens in a new tab) is a curated and repaired subset of BIRD’s training data. The authors audited roughly 2,500 examples and found errors in 61.1%, including incorrect reference SQL in more than half. They repaired the questions, supporting evidence, and reference queries to give training a more reliable reward signal.
The verified training data had the clearest measurable effect on performance. We trained Kimi K2.6, GLM-5.3 Flash, and Qwen3.8-27B with the same algorithm and questions on two versions of the annotations: the paper’s cleaned version and BIRD’s original. Across the first two epochs, the cleaned annotations scored 9.4 points higher on Arcwise-Plat on average.
With the original annotations, performance actually degraded over training, with all three models falling 7 to 10 points from their early peaks. For Kimi, the advantage also carried to Spider2-Lite, where the cleaned data scored 10.4 points higher.
Cleaned vs. Original Annotations
Arcwise-Plat-SQL accuracy over the first two epochs
We noticed that validation accuracy still seemed to plateau across runs in the 92–93% range, which suggested there may still be issues in the BIRD-Platinum dataset. We investigated and found there was still a considerable number of inconsistencies and errors remaining.
Residual Issues in BIRD-Platinum
Each question counted once · 2,462 questions
- Incorrect Gold SQL159 · 6.5%
- Incorrect Evidence220 · 8.9%
- Ambiguity in the Question337 · 13.7%
- Missing Evidence137 · 5.6%
- Evidence Is a Hint89 · 3.6%
- Readability21 · 0.9%
- Other7 · 0.3%
- Unchanged1,492 · 60.6%
Categories are ranked by severity, many questions fell under multiple categories.
Multiple issues per question · 2,462 questions
A question can appear in several categories, so the bars add up to more than the 39.4% of questions (970) with any issue. Each bar uses all 2,462 questions as its denominator.
What's compelling about this categorization is, the 6.5% bucket of Incorrect Gold SQL is about the same gap as the validation accuracy plateau when training on BIRD-Platinum.
The BIRD-Diamond Dataset
We invested some considerable effort into building BIRD-Diamond (opens in a new tab), a further cleaned version built on top of the authors’ BIRD-Platinum work. We were optimistic that we would see more gains from this additional dataset cleaning, but surprisingly, we actually saw an average of about 1% lower eval scores when trained on the Diamond dataset.
We dug into our runs to understand why this second dataset cleaning harmed our eval strength. Our first clue was in the validation scores.
Brackets show each run's gain from step 0 to step 65. BIRD-Diamond starts highest but gains the least: 6.6 points, against 10.6 for BIRD-Platinum and 10.7 for BIRD.
While fixing the incorrect gold SQL closed the gap for the models to saturate the training set, the disambiguation fixes removed some of the complexity of the dataset. We can see above that the learnable delta is essentially the same between BIRD and BIRD-Platinum. For BIRD-Diamond, however, it appears we trivialized about 4% of the problems through cleaning issues beyond incorrect gold SQL queries.
The original datasets teach the model to be more robust to such defects in the questions. We found that the 1% delta on eval from BIRD-Diamond runs vs BIRD-Platinum was dominated by a specific class of question defects: evidence errors. Training on BIRD-Diamond teaches the model to be extremely trusting of the evidence, while training on BIRD-Platinum teaches some skepticism and robustness to handle such errors in the Arcwise eval set.
Reward Shaping
Reward shaping adds smaller signals on top of the outcome reward to steer how the model reaches its answer. ReViSQL adds two.
VeriEQL Penalty
Targets queries that are right by accident.
If an equivalence checker, VeriEQL, proves that a matching query would return a different result from the gold query on some other database, the reward drops.
- ✓ Correct0.8
- × Incorrect0
- – No solution−1
Evidence Process Reward
Targets answers that ignore the evidence.
The model must restate the prompted evidence early in its response, then check its final answer against it. Each skipped step costs 0.1.
- ✓ Correct0.9
- × Incorrect−0.1
- – No solution−1.1
VeriEQL Penalty
VeriEQL (opens in a new tab) asks a stricter question than the benchmark grader: not just whether two queries match on the real database, but whether they would return the same result on any valid database. It searches for a small counterexample where they disagree. If it finds one, ReViSQL treats the model’s answer as accidentally correct. In practice, that also makes VeriEQL a kind of style reward, favoring queries that follow the same conventions as BIRD’s gold SQL.
VeriEQL produced counterexamples for 29% of the model’s benchmark-correct answers, close to the paper’s reported 32.8%. But most did not survive our benchmark-aligned check. Put differently, for every 100 answers the benchmark marked correct, VeriEQL found a counterexample for about 29, but only about 5 still produced different results under the benchmark’s own grader.
We also tested three other SQL equivalence checkers, SpotIt+, ParSEval, and Polygon, and none was a clear improvement over VeriEQL. The investigation surfaced a paradox: the stricter a checker’s definition of “the same result,” the more often it rejected benchmark-correct answers over technical differences the benchmark itself does not count. In other words, a more aggressive checker does not necessarily produce a better reward signal; it can just add more noise.
Evidence Process Reward
The evidence process reward behaved differently. Models learned the required format quickly: fewer than 1% of rollouts skipped the evidence steps after the first epoch. That shows the reward successfully teaches compliance. We suspect it also interacts with the eval questions whose evidence is misleading: drawing more attention to the evidence may encourage the model to question it rather than follow it blindly.
Measuring the Effect
In the paper's ablation, removing reward shaping lowered Arcwise-Plat accuracy by 2.8 percentage points. In our own ablation runs on GLM-5.3 Flash, the run without the VeriEQL penalty peaked 0.8 percentage points lower than its shaped counterpart, and the run without the evidence reward 1.4 points lower.
The chart below follows both pairs of runs, showing each run’s accuracy at every checkpoint.
Arcwise-Plat-SQL accuracy at each checkpoint
VeriEQL penalty
Evidence reward
Each panel compares a run trained with that shaping signal against its counterpart trained without it, evaluated every 5 steps with our stricter grader.
Shaping appears to be beneficial in general, and most pronounced in the 2–3 epoch window, where runs tend to start to plateau. Additionally, the evidence reward appears to have a stabilizing effect on the runs, which is interesting.
We will revisit shaping again below, as it had some interesting interactions with some of the recipe modifications we tested.
Reasoning Settings
We expected higher reasoning settings and more tokens per turn to improve performance, but the effect was more nuanced. DeepSeek V4.1 Flash benefited from its highest setting with longer turns: at max effort and 8,192 tokens per turn, it reached its peak in about half the steps it needed at 3,072. Qwen3.8 showed no difference between medium and xhigh at 8,192 tokens per turn. Our recipe defaults to xhigh to be safe, but the setting does not seem to have a strong effect on performance.
Another surprising result was that GLM-5.3 Flash did slightly better at 3,072 tokens per turn than at either longer budget we tested, and so did Nemotron 3.5 Lightning. This pattern does not hold for every model, so these settings should be tested and tuned for any new model added to the River API.
Extending the Recipe
Once we could reliably reproduce the paper’s results, we began looking for ways to extend the recipe and push performance higher.
Training Hints
During training, when all 16 attempts at a question fail, the step gets no learning signal from that question. We added a fallback to try to recover one.
We re-sample the question with a hint added to the prompt and use those results instead. We pre-generated hints (opens in a new tab) with Kimi K2.6 using the question, its evidence, the schema, and the reference query.
After the first epoch of our final runs, about 6% of questions needed a hint, and surprisingly, that share barely varied across models.
While that share was similar across models, the response to hints was not. After receiving a hint, 85% of hinted re-runs came back 16/16 correct on GLM-5.3 Flash. Yet for Nemotron 3.5 Lightning, only 43% of hinted re-runs came back 16/16 correct.
Ideally with hints, you want the model to only recover between 1 and 15 correct attempts on a question so that there is signal to learn from. So hints give stronger models less to learn from: they often solve the hinted question outright.
Hints were one of the strongest gains we found from changing the recipe. The run with hints stayed ahead for most of training, leading at 26 of the 33 evaluations both runs have reached, and its best score over those steps was about a point and a half higher, 92.8% against 91.2%.
Hints matter because they are the first part of the recipe that can teach the model something beyond what it already knows. Reinforcement learning in the ReViSQL recipe works by making the model more likely to choose correct answers it can already produce, so a question it never answers correctly gives it nothing to learn from. A hint lets the model reach a correct answer on those questions, and what it learned there carried over to evaluation, where hints are never shown.
Our hints were generated by a model with access to the question, evidence, and gold SQL. In practice, many appear to be too strong, making the solution too easy to recover. A better hint set, ideally human-written and limited to smaller, more targeted nudges, could make this fallback more effective.
Learnability Weighting
Most training questions stop teaching the model anything early on. About 60% of questions were already solved in all 16 attempts on their first visit. With a standard curriculum, the share of questions with signal falls steadily as the model solves more of the training set, so later steps spend most of their compute on questions that teach nothing. Learnability weighting counters this by steering training toward the questions that still have something to teach.
Learnability weighting only takes effect starting at epoch 2. Higher percentage of partly solved questions means more to learn per step at that time.
Shows that the average reward signal per question decreases as the pool becomes more solved.
We enable learnability weighting at experimentally-tuned values for our recommended recipes.
BIRD-Platinum-Hard
After the first epoch, most questions come back 16/16 on every visit. A natural next idea is to drop them altogether and save the inference. We tried this two ways: training each model only on the questions it had not solved 16/16 on its first visit, and training on a shared set we call BIRD-Platinum-Hard (opens in a new tab), the 689 training questions that most models could not solve fully on the first visit. A pass over BIRD-Platinum-Hard costs a third of the questions sampled per epoch.
Behavior on this dataset was different. Performance rose much faster early in training, then plateaued sooner. It was also less stable for models like Qwen 3.6, as removing easier questions weakened the baseline that helps keep the model from drifting too far.
The reduced exposure to fully solved questions also weakened the effect of the VeriEQL shaping reward. Under the base reward, these questions provide no learning signal because every attempt receives the same score. Shaping reintroduces differences within them, giving the optimizer useful signal. Because the Hard dataset removes most of these solved questions, there are fewer opportunities for shaping to help, and peak evaluation performance was lower with shaping enabled.
This result also suggests that learnability weighting should be careful not to suppress the solved cohort so aggressively that it removes useful shaping signal.
Recipe Cards
Every model result in this post was trained on the same core recipe. Its preset changes only a handful of settings, and everything else is shared. Pick a model to see its card.
Kimi K2.6
nvidia/Kimi-K2.6-NVFP4
preset kimi-k2.6
| Reasoning effort | Model default |
|---|---|
| Tokens per turn | 3,072 |
| Training length | 5 epochs · 165 steps |
| Tool-format penalty | Off |
| Shared by every preset | |
| Training data | BIRD-Platinum (opens in a new tab), 2,064 train / 398 validation |
| RL objective | CISPO |
| Expert routing replay | On1 |
| LoRA rank | 32 |
| Learning rate | 5 × 10−5, 10-step warmup, cosine decay to 10% |
| Batch size | 64 questions × 16 attempts |
| SQL tool calls | Up to 5, then a final answer |
| Shaping | VeriEQL penalty 0.2; evidence steps −0.1 each |
| Curriculum | Learnability weighting from epoch 2; training hints on |
Nemotron 3.5 Lightning
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
preset nemotron-3.5-lightning
| Reasoning effort | Model default |
|---|---|
| Tokens per turn | 3,072 |
| Training length | 7 epochs · 231 steps |
| Tool-format penalty | Off |
| Shared by every preset | |
| Training data | BIRD-Platinum (opens in a new tab), 2,064 train / 398 validation |
| RL objective | CISPO |
| Expert routing replay | On1 |
| LoRA rank | 32 |
| Learning rate | 5 × 10−5, 10-step warmup, cosine decay to 10% |
| Batch size | 64 questions × 16 attempts |
| SQL tool calls | Up to 5, then a final answer |
| Shaping | VeriEQL penalty 0.2; evidence steps −0.1 each |
| Curriculum | Learnability weighting from epoch 2; training hints on |
GLM-5.3 Flash
zai-org/GLM-5.3-Flash
preset glm-5.3-flash
| Reasoning effort | Max |
|---|---|
| Tokens per turn | 3,072 |
| Training length | 7 epochs · 231 steps |
| Tool-format penalty | Off |
| Shared by every preset | |
| Training data | BIRD-Platinum (opens in a new tab), 2,064 train / 398 validation |
| RL objective | CISPO |
| Expert routing replay | On1 |
| LoRA rank | 32 |
| Learning rate | 5 × 10−5, 10-step warmup, cosine decay to 10% |
| Batch size | 64 questions × 16 attempts |
| SQL tool calls | Up to 5, then a final answer |
| Shaping | VeriEQL penalty 0.2; evidence steps −0.1 each |
| Curriculum | Learnability weighting from epoch 2; training hints on |
DeepSeek V4.1 Flash
deepseek-ai/DeepSeek-V4.1-Flash
preset deepseek-v4.1-flash
| Reasoning effort | Max |
|---|---|
| Tokens per turn | 8,192 |
| Training length | 5 epochs · 165 steps |
| Tool-format penalty | Off |
| Shared by every preset | |
| Training data | BIRD-Platinum (opens in a new tab), 2,064 train / 398 validation |
| RL objective | CISPO |
| Expert routing replay | On1 |
| LoRA rank | 32 |
| Learning rate | 5 × 10−5, 10-step warmup, cosine decay to 10% |
| Batch size | 64 questions × 16 attempts |
| SQL tool calls | Up to 5, then a final answer |
| Shaping | VeriEQL penalty 0.2; evidence steps −0.1 each |
| Curriculum | Learnability weighting from epoch 2; training hints on |
Qwen3.8-27B
Qwen/Qwen3.8-27B-FP8
preset qwen3.8-27b
| Reasoning effort | xhigh |
|---|---|
| Tokens per turn | 8,192 |
| Training length | 5 epochs · 165 steps |
| Tool-format penalty | 0.05 per malformed call, up to 0.2 |
| Shared by every preset | |
| Training data | BIRD-Platinum (opens in a new tab), 2,064 train / 398 validation |
| RL objective | CISPO |
| Expert routing replay | On1 |
| LoRA rank | 32 |
| Learning rate | 5 × 10−5, 10-step warmup, cosine decay to 10% |
| Batch size | 64 questions × 16 attempts |
| SQL tool calls | Up to 5, then a final answer |
| Shaping | VeriEQL penalty 0.2; evidence steps −0.1 each |
| Curriculum | Learnability weighting from epoch 2; training hints on |
Qwen3.6-35B-A3B
Qwen/Qwen3.6-35B-A3B-FP8
preset qwen3.6-35b-a3b
| Reasoning effort | Model default |
|---|---|
| Tokens per turn | 3,072 |
| Training length | 5 epochs · 165 steps |
| Tool-format penalty | 0.05 per malformed call, up to 0.2 |
| Shared by every preset | |
| Training data | BIRD-Platinum (opens in a new tab), 2,064 train / 398 validation |
| RL objective | CISPO |
| Expert routing replay | On1 |
| LoRA rank | 32 |
| Learning rate | 5 × 10−5, 10-step warmup, cosine decay to 10% |
| Batch size | 64 questions × 16 attempts |
| SQL tool calls | Up to 5, then a final answer |
| Shaping | VeriEQL penalty 0.2; evidence steps −0.1 each |
| Curriculum | Learnability weighting from epoch 2; training hints on |
1 Expert routing is only available on MoE models
Reflection on Results
One of the most exciting results was how consistently the recipe worked across models. With only small per-model adjustments, the same recipe improved six models from five families by 9 to 36 points on Arcwise-Plat. What surprised us most is where that gain comes from. A training group only teaches the model something when some of its 16 attempts succeed, so most of the improvement comes from reinforcing answers the models could already produce, even if only rarely, until they produce them reliably.
However, this left us with a question: if training mostly sharpens what each model can already do, why did models that started far apart all end up in the same place? Five of the six models finished within about a point and a half of each other on Arcwise-Plat, between 92.0% and 93.4%, even though their base models started anywhere from 71% to 83%. On Spider2-Lite, the same five models stayed spread out, finishing between 50.4% and 61.5%. That contrast made us ask whether Arcwise-Plat itself had become the limit, so we looked at how often each question was solved across 979 evaluations of those five models.
For each question, we checked how often our strongest runs solved it in evaluations from step 60 on. The first 438 questions, all solved at least 93% of the time, are compressed on the left.
The chart shows, for every question, how often the model that handles it best gets it right in evaluations from step 60 on. Nearly 90% of the questions are solved reliably, in at least 90% of evaluations. The next 7% of questions seem to be reachable, but much less reliably, with the models capturing only a fraction of them on any given checkpoint. While around 3% of the benchmark looks out of reach for any recipe.
One particular class of questions seems to explain a lot of the behavior we saw late in the run.
Ambiguous Questions
Evaluation scores tend to move in small increments, but those shifts hide much more activity underneath. Between evaluations five steps apart, 26 questions change outcome on average. Most changes cancel out as answers flip in both directions, so the score reflects only the net difference.
Much of that churn comes from a recurring pool of questions rather than being spread evenly across the benchmark.
When we investigated these questions, we found that most had ambiguity with multiple valid interpretations. When the model learns to favor one interpretation, it unlocks some questions, while losing others. This also explains why some checkpoint evals are stronger than others. The models settled on more net-positive interpretations at those steps.
“Please provide the full name of the away team that scored the most goals.”
Most goals in one match, or across a season?
“What is the highest monthly consumption in the year 2012?”
One customer’s month, or the month’s total?
“Calculate the average number of oxygen atoms in single-bonded molecules.”
At least one single bond, or only single bonds?
If the same ambiguity is interpreted differently across questions, no single reading can match them all. Some portion of the benchmark is therefore unreachable, even for a perfect model.
Final Thoughts
A stronger model might still push Arcwise-Plat a little higher, but we are probably close to where this benchmark saturates. Spider2-Lite, on the other hand, is still far from saturated, and our trained models rank roughly in the same order as their base models did.
The clearest gain in this recipe came from BIRD-Platinum’s cleaned annotations. But most of those questions are already easily solved by the strongest open-weight models. That leaves only a shrinking minority of the dataset providing useful training signal. Pushing further on harder benchmarks like Spider2-Lite will likely require harder training data annotated to the same standard.
Getting Started
We encourage you to try this recipe on your own data.
Adding your own data is straightforward. Each example needs a question over one of your databases paired with the correct SQL query. You can often bootstrap this data from existing production query logs, then write the corresponding questions afterward.
Evidence is optional, and we generally recommend leaving it out when adapting the recipe to your own data. BIRD-Platinum and BIRD-Diamond become harder without evidence, and that setting is also closer to how real users behave. They rarely specify the exact columns, joins, or formulas needed to answer a question.
If you want more training volume, you can also mix in BIRD-Platinum (opens in a new tab) or BIRD-Diamond (opens in a new tab). When using BIRD data, we recommend disabling shaping or applying it only to your own examples so the model is shaped toward the style you actually care about.
If you want to build on this recipe and would like a hand, contact us.
References
Yuxuan Zhu, Tengjun Jin, Yoojin Choi, and Daniel Kang. (2026). Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering. arXiv:2603.20004v4.
Paper (opens in a new tab) · ReViSQL code (opens in a new tab)