Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models
Quick summary
arXiv:2605.25765v2 Announce Type: replace-cross Abstract: Existing closed-form methods for concept unlearning in text-to-image diffusion models typically derive editing directions from fixed text embeddings, which may not fully capture how concepts are expressed across latent states, timesteps, and layers. To capture this variation, we investigate cross-attention activations collected during denoising. In controlled probing experiments using the same anchor prompts, activation-derived bases achieve approximately five times the recall of text-derived bases on held-out prompts expressing the tar
Key takeaways
- arXiv:2605.25765v2 Announce Type: replace-cross Abstract: Existing closed-form methods for concept unlearning in text-to-image diffusion models typically derive editing directions from fixed text embeddings, which may not fully capture how concepts are expressed across latent states, timesteps, and layers.
- To capture this variation, we investigate cross-attention activations collected during denoising.
- In controlled probing experiments using the same anchor prompts, activation-derived bases achieve approximately five times the recall of text-derived bases on held-out prompts expressing the tar
Why it matters
“Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments