arXiv Artificial Intelligence

Hidden not Deleted: How Networks Suppress Entangled Features

Hidden not Deleted: How Networks Suppress Entangled Features

Quick summary

arXiv:2609.27593v1 Announce Type: cross Abstract: Concept erasure methods that operate via linear projection assume that features occupy separable subspaces. We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target. Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions. We map

Key takeaways

  • arXiv:2609.27593v1 Announce Type: cross Abstract: Concept erasure methods that operate via linear projection assume that features occupy separable subspaces.
  • We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target.
  • Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions.

Why it matters

“Hidden not Deleted: How Networks Suppress Entangled Features” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗