FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation
Quick summary
arXiv:2609.36670v1 Announce Type: new Abstract: A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propa
Key takeaways
- arXiv:2609.36670v1 Announce Type: new Abstract: A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable.
- Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization.
- While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propa
Why it matters
The importance of “FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments