Exploring Bottom-Up Clustering for Creating Semantic IDs
Quick summary
arXiv:2609.08310v1 Announce Type: cross Abstract: The success of generative retrieval has largely been attributed to the use of Semantic IDs, which improve over arbitrary item-level identifiers such as hashes by capturing the semantics of items. The main challenges faced when constructing Semantic IDs, however, is in mapping each identifier to a unique product and capturing information valuable to downstream tasks. Past works have appended additional codewords to de-duplicate item identifiers and utilized residual quantization to create hierarchical clusters. In this work, we present an algori
Key takeaways
- arXiv:2609.08310v1 Announce Type: cross Abstract: The success of generative retrieval has largely been attributed to the use of Semantic IDs, which improve over arbitrary item-level identifiers such as hashes by capturing the semantics of items.
- The main challenges faced when constructing Semantic IDs, however, is in mapping each identifier to a unique product and capturing information valuable to downstream tasks.
- Past works have appended additional codewords to de-duplicate item identifiers and utilized residual quantization to create hierarchical clusters.
Why it matters
“Exploring Bottom-Up Clustering for Creating Semantic IDs” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments