arXiv Artificial Intelligence

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

Quick summary

arXiv:2609.37798v1 Announce Type: cross Abstract: Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-g

Key takeaways

  • arXiv:2609.37798v1 Announce Type: cross Abstract: Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks.
  • However, current approaches rely on complex, carefully designed prediction targets.
  • We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders.

Why it matters

“GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗