FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
Quick summary
arXiv:2609.03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode. We introduce FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses. For each on-policy trajectory, a frozen training-time view of the same policy uses privileged context to produce token-level log-probability g
Key takeaways
- arXiv:2609.03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.
- We introduce FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses.
- For each on-policy trajectory, a frozen training-time view of the same policy uses privileged context to produce token-level log-probability g
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience” may reshape data collection, model training, output accountability and market access.

Member comments