arXiv Artificial Intelligence

Asymmetric Within-Document Predictive Learning for Scientific Document Representation

Asymmetric Within-Document Predictive Learning for Scientific Document Representation

Quick summary

arXiv:2608.28625v1 Announce Type: cross Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations. Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using

Key takeaways

  • arXiv:2608.28625v1 Announce Type: cross Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers.
  • We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations.
  • Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using

Why it matters

“Asymmetric Within-Document Predictive Learning for Scientific Document Representation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗