arXiv Artificial Intelligence

Pre-train to Gain: Robust Learning Without Clean Labels

Pre-train to Gain: Robust Learning Without Clean Labels

Quick summary

arXiv:2511.20844v2 Announce Type: replace-cross Abstract: Training deep networks with noisy labels leads to poor generalization and degraded accuracy due to overfitting to label noise. Existing approaches for learning with noisy labels often rely on the availability of a clean subset of data. By pre-training a feature extractor on the target dataset without labels using in-domain self-supervised learning (SSL), followed by standard supervised training on the same noisy dataset, we can train a more noise robust model without requiring a subset with clean labels. We evaluate both contrastive and

Key takeaways

  • arXiv:2511.20844v2 Announce Type: replace-cross Abstract: Training deep networks with noisy labels leads to poor generalization and degraded accuracy due to overfitting to label noise.
  • Existing approaches for learning with noisy labels often rely on the availability of a clean subset of data.
  • By pre-training a feature extractor on the target dataset without labels using in-domain self-supervised learning (SSL), followed by standard supervised training on the same noisy dataset, we can train a more noise robust model without requiring a subset with clean labels.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗