arXiv Artificial Intelligence

PUFFER: Incremental Fuzzy Deduplication for Continuously Evolving Corpora

PUFFER: Incremental Fuzzy Deduplication for Continuously Evolving Corpora

Quick summary

arXiv:2608.28622v1 Announce Type: cross Abstract: Large language model training corpora grow through successive, often redundant releases, so each release must be deduplicated against both itself and the accumulated history. At trillion-token scale, this requires incremental ingestion, bounded resident memory, deterministic retry, and dataset-scoped lifecycle control without repeated corpus-wide rebuilding. We introduce PUFFER (Provenance-aware Updatable Fuzzy Filtering for Evolving Repositories), a MinHash-LSH fuzzy-deduplication pipeline built around two design choices. First, PUFFER stores

Key takeaways

  • arXiv:2608.28622v1 Announce Type: cross Abstract: Large language model training corpora grow through successive, often redundant releases, so each release must be deduplicated against both itself and the accumulated history.
  • At trillion-token scale, this requires incremental ingestion, bounded resident memory, deterministic retry, and dataset-scoped lifecycle control without repeated corpus-wide rebuilding.
  • We introduce PUFFER (Provenance-aware Updatable Fuzzy Filtering for Evolving Repositories), a MinHash-LSH fuzzy-deduplication pipeline built around two design choices.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗