Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
Quick summary
arXiv:2610.01378v1 Announce Type: new Abstract: Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit thi
Key takeaways
- arXiv:2610.01378v1 Announce Type: new Abstract: Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused.
- A waveform-label pair does not preserve this knowledge.
- We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity.
Why it matters
“Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments