科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-09-13· cs.LG

Learning Source Acquisition Policies by Offline Planning

Ziqi Zhao, Run Xu, Qingjian Ni

原始摘要(英文原文)· Original abstract
Predicting under an acquisition budget requires choosing feature groups whose value can depend on later queries. O-MPAC transfers finite-horizon risk-cost targets from complete training records into a shared source-action scorer. At inference time, the scorer uses partial observations and source metadata, re-scores after each query, and applies a hard cost mask. We analyze how tied teacher targets and the remaining planning horizon affect the learned decisions. Uniform supervision over tied minima preserves the target distribution under source relabeling. In a five-seed routing experiment, it achieves 0.965 accuracy under both original and context-last orders. On six real tasks, validation selects H1 without action cross-entropy in all thirty splits. O-MPAC has the highest mean budget-integrated accuracy on five tasks against source-adapted GDFS, DIME, AACO+NN and a static policy.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Learning Source Acquisition Policies by Offline Planning — 科研速览 Science Skim