arXiv Artificial IntelligenceMitigating Reasoning-Induced Misalignment via Safety-Direction Penalty
arXiv:2608.23497v1 Announce Type: new Abstract: Reasoning-Induced Misalignment, where fine-tuning on…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2608.23497v1 Announce Type: new Abstract: Reasoning-Induced Misalignment, where fine-tuning on…
arXiv Artificial IntelligencearXiv:2608.23493v1 Announce Type: new Abstract: Self-reflection is a powerful mechanism for credit assignment…
arXiv Artificial IntelligencearXiv:2608.23484v1 Announce Type: new Abstract: We present Team Semiintelligencn's solution for the ACM…
arXiv Artificial IntelligencearXiv:2608.23475v1 Announce Type: new Abstract: As large language models are increasingly used in data-scarce…
arXiv Artificial IntelligencearXiv:2608.23417v1 Announce Type: new Abstract: Agent skills are reusable procedural artifacts that extend…
arXiv Artificial IntelligencearXiv:2608.23373v1 Announce Type: new Abstract: Forecasting long-range influenza-like illness (ILI) matters…
arXiv Artificial IntelligencearXiv:2608.23318v1 Announce Type: new Abstract: Hint-based reinforcement learning addresses reward sparsity…
arXiv Artificial IntelligencearXiv:2608.23313v1 Announce Type: new Abstract: Vision-language model safety benchmarks typically evaluate…
arXiv Artificial IntelligencearXiv:2608.23264v1 Announce Type: new Abstract: Although Large Language Models (LLMs) are aligned to optimize…
arXiv Artificial IntelligencearXiv:2608.23263v1 Announce Type: new Abstract: The FAIR Digital Object (FDO) framework mandates that…
arXiv Artificial IntelligencearXiv:2608.23256v1 Announce Type: new Abstract: Recent work proposes next-chunk reasoning RL for leveraging…
arXiv Artificial IntelligencearXiv:2608.23218v1 Announce Type: new Abstract: Advances in neural theorem provers have been impressive, but…