On the role of the tokenizer in ECG transformer models
Quick summary
arXiv:2609.15433v1 Announce Type: cross Abstract: Tokenization determines both the physiological content presented to an ECG Transformer and the sequence over which attention operates. We compare eight tokenization strategies across Transformer, Informer, Reformer, and FEDformer on the nine-label CPSC2018 classification task. The input projection and principal backbone capacity are controlled to isolate the effect of token construction. Median-beat and HeartLang tokenization achieve mean macro-AUCs of 0.893 and 0.889 across the four backbones, compared with 0.822 and 0.824 for point-wise and p
Key takeaways
- arXiv:2609.15433v1 Announce Type: cross Abstract: Tokenization determines both the physiological content presented to an ECG Transformer and the sequence over which attention operates.
- We compare eight tokenization strategies across Transformer, Informer, Reformer, and FEDformer on the nine-label CPSC2018 classification task.
- The input projection and principal backbone capacity are controlled to isolate the effect of token construction.
Why it matters
“On the role of the tokenizer in ECG transformer models” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments