TTSR: Test-Time Self-Evolving via Reflection
Quick summary
arXiv:2603.03297v3 Announce Type: replace-cross Abstract: Test-time training (TTT) adapts large language models (LLMs) during inference using only unlabeled test inputs. Existing methods, however, face two major bottlenecks on hard reasoning tasks: (1) \emph{lack of learnable samples}, as self-generated pseudo-labels on difficult questions are often noisy and yield unstable rewards; and (2) \emph{inefficient exploration}, as performance gains depend on repeatedly sampling many rollouts without explicit diagnosis of why previous attempts fail. We propose \textbf{TTSR} (\textbf{T}est-\textbf{T}i
Key takeaways
- arXiv:2603.03297v3 Announce Type: replace-cross Abstract: Test-time training (TTT) adapts large language models (LLMs) during inference using only unlabeled test inputs.
- Existing methods, however, face two major bottlenecks on hard reasoning tasks: (1) \emph{lack of learnable samples}, as self-generated pseudo-labels on difficult questions are often noisy and yield unstable rewards; and (2) \emph{inefficient exploration}, as performance gains depend on repeatedly sampling many rollouts without explicit diagnosis of why previous attempts fail.
- We propose \textbf{TTSR} (\textbf{T}est-\textbf{T}i
Why it matters
“TTSR: Test-Time Self-Evolving via Reflection” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments