arXiv Artificial Intelligence

How Do LLMs Change Predictions Under Negation?

How Do LLMs Change Predictions Under Negation?

Quick summary

arXiv:2610.09571v2 Announce Type: replace-cross Abstract: Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?"). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by

Key takeaways

  • arXiv:2610.09571v2 Announce Type: replace-cross Abstract: Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it.
  • We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?").
  • To understand and address this brittleness, we mechanistically examine how models operate under negation.

Why it matters

“How Do LLMs Change Predictions Under Negation?” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗