arXiv Artificial Intelligence

CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

Quick summary

arXiv:2508.08879v3 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures. Prior work has evaluated cultural awareness in LLMs by analysing their outputs. This approach overlooks how cultures are represented within the model parameters, missing why models generate incorrect responses. To bridge this gap, we propose CulTrace, a mechanistic interpretability-based method that probes the internal representations of LLMs for cultural kno

Key takeaways

  • arXiv:2508.08879v3 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures.
  • Prior work has evaluated cultural awareness in LLMs by analysing their outputs.
  • This approach overlooks how cultures are represented within the model parameters, missing why models generate incorrect responses.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗