arXiv Artificial IntelligenceRefusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
arXiv:2605.28553v2 Announce Type: replace Abstract: In this paper, we investigate whether refusal behavior…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2605.28553v2 Announce Type: replace Abstract: In this paper, we investigate whether refusal behavior…
arXiv Artificial IntelligencearXiv:2605.28025v2 Announce Type: replace Abstract: Existing safety evaluations for large language models…
arXiv Artificial IntelligencearXiv:2605.12462v2 Announce Type: replace Abstract: Extreme weather and volatile wholesale electricity…
arXiv Artificial IntelligencearXiv:2604.01658v3 Announce Type: replace Abstract: Large language model (LLM)-based evolution is a promising…
arXiv Artificial IntelligencearXiv:2602.10009v3 Announce Type: replace Abstract: Large Language Models (LLMs) are unable to reliably…
arXiv Artificial IntelligencearXiv:2602.09341v2 Announce Type: replace Abstract: Multi-agent systems (MAS) can substantially extend the…
arXiv Artificial IntelligencearXiv:2602.01207v2 Announce Type: replace Abstract: Offline preference optimization aligns reasoning models…
arXiv Artificial IntelligencearXiv:2602.00266v2 Announce Type: replace Abstract: Two deep ReLU networks can have entirely different…
arXiv Artificial IntelligencearXiv:2601.16027v3 Announce Type: replace Abstract: The rise of live streaming has transformed online…
arXiv Artificial IntelligencearXiv:2601.10029v3 Announce Type: replace Abstract: Academic paper search is a fundamental task in scientific…
arXiv Artificial IntelligencearXiv:2510.15221v3 Announce Type: replace Abstract: Affective computing has matured rapidly in laboratory…
arXiv Artificial IntelligencearXiv:2505.19030v5 Announce Type: replace Abstract: Large language models (LLMs) are increasingly expected to…