TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization
Quick summary
arXiv:2609.39033v1 Announce Type: cross Abstract: CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions. The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-norma
Key takeaways
- arXiv:2609.39033v1 Announce Type: cross Abstract: CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity.
- We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions.
- The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-norma
Why it matters
“TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments