POLARIS: Guiding Small Models to Write Long Stories
Quick summary
arXiv:2606.04095v2 Announce Type: replace-cross Abstract: Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quality significantly degrades as length increases, especially when compared to frontier models. We present POLARIS (Policy Optimization with LLM-as-a-judge rewards and Anchored-Reference Injection for Storywriting), a lower-compute GRPO recipe with two key ingredients: a frontier LLM judge with a structured Story Quality rubric as the online reward, and human-reference injection (HRI), where a
Key takeaways
- arXiv:2606.04095v2 Announce Type: replace-cross Abstract: Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quality significantly degrades as length increases, especially when compared to frontier models.
- We present POLARIS (Policy Optimization with LLM-as-a-judge rewards and Anchored-Reference Injection for Storywriting), a lower-compute GRPO recipe with two key ingredients: a frontier LLM judge with a structured Story Quality rubric as the online reward, and human-reference injection (HRI), where a
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “POLARIS: Guiding Small Models to Write Long Stories” may reshape data collection, model training, output accountability and market access.

Member comments