arXiv Artificial IntelligenceInference-Time Nash Alignment
arXiv:2609.08082v1 Announce Type: new Abstract: Preference-based fine-tuning methods such as RLHF and DPO…
Curated from international AI laboratories, specialist publications and technology outlets. Last update: 3 hours ago.
arXiv Artificial IntelligencearXiv:2609.08082v1 Announce Type: new Abstract: Preference-based fine-tuning methods such as RLHF and DPO…
arXiv Artificial IntelligencearXiv:2609.08071v1 Announce Type: new Abstract: Firms making inventory decisions have access to operational…
arXiv Artificial IntelligencearXiv:2609.08062v1 Announce Type: new Abstract: Tool-using language agents can delegate and revoke…
arXiv Artificial IntelligencearXiv:2609.08025v1 Announce Type: new Abstract: Reasoning agents increasingly rely on external tools such as…
arXiv Artificial IntelligencearXiv:2609.08016v1 Announce Type: new Abstract: Multi-agent debate, in which several LLMs exchange arguments…
arXiv Artificial IntelligencearXiv:2609.08015v1 Announce Type: new Abstract: Long-running AI agents may read state, reason, wait for tools…
arXiv Artificial IntelligencearXiv:2609.08003v1 Announce Type: new Abstract: Behavioral foundation models have been proposed as stand-ins…
arXiv Artificial IntelligencearXiv:2609.07998v1 Announce Type: new Abstract: We study the control of Markov decision processes in which…
arXiv Artificial IntelligencearXiv:2609.07987v1 Announce Type: new Abstract: LLM-based digital twins promise to reduce repeated human data…
arXiv Artificial IntelligencearXiv:2609.07984v1 Announce Type: new Abstract: Process mining has long turned event logs into process…
arXiv Artificial IntelligencearXiv:2609.07954v1 Announce Type: new Abstract: Sparse Sinkhorn layers use a fixed support graph to restrict…
arXiv Artificial IntelligencearXiv:2609.07944v1 Announce Type: new Abstract: Existing causal-inference benchmarks for LLMs mostly score…