arXiv Artificial Intelligence

Causal and Interpretable Structures in LLM Compositional Tasks

Causal and Interpretable Structures in LLM Compositional Tasks

Quick summary

arXiv:2609.35970v1 Announce Type: cross Abstract: Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layer

Key takeaways

  • arXiv:2609.35970v1 Announce Type: cross Abstract: Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them.
  • How is such relational information represented and processed across transformer layers?
  • We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗