Hidden Gauge Controls Feature Specialization in ReLU Networks
Quick summary
arXiv:2608.06766v2 Announce Type: replace-cross Abstract: The success of deep learning depends on learning useful representations, yet predicting how training organizes these representations across neurons remains difficult. In this work, we show that changing the scale of initial weights can determine which neurons learn a feature without altering any neuron's initial contribution. We construct ReLU networks with identical initial features and predictions that reach the same final predictions with different roles for their neurons. In one, all neurons share the learned feature. In the other,
Key takeaways
- arXiv:2608.06766v2 Announce Type: replace-cross Abstract: The success of deep learning depends on learning useful representations, yet predicting how training organizes these representations across neurons remains difficult.
- In this work, we show that changing the scale of initial weights can determine which neurons learn a feature without altering any neuron's initial contribution.
- We construct ReLU networks with identical initial features and predictions that reach the same final predictions with different roles for their neurons.
Why it matters
“Hidden Gauge Controls Feature Specialization in ReLU Networks” is a product decision that may change how people work with AI. Its value depends on task completion, correction effort and data handling—not simply the presence of a new feature.

Member comments