arXiv Artificial Intelligence

Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation

Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation

Quick summary

arXiv:2610.10974v1 Announce Type: cross Abstract: Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over con

Key takeaways

  • arXiv:2610.10974v1 Announce Type: cross Abstract: Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation.
  • Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy.
  • We study budgeted acquisition of such annotations for contextual-bandit OPE.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗