Asymmetric Within-Document Predictive Learning for Scientific Document Representation
Quick summary
arXiv:2608.28625v1 Announce Type: cross Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations. Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using
Key takeaways
- arXiv:2608.28625v1 Announce Type: cross Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers.
- We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations.
- Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using
Why it matters
“Asymmetric Within-Document Predictive Learning for Scientific Document Representation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments