科研速览 · Science Skim继续刷下去 · Keep skimming →
2026-09-11· Reinforcement learning

DYNAMIC SYMBOLIC DEEP REINFORCEMENT LEARNING STATE-DIFFERENTIATED REWARDS AND ADAPTIVE WORLD MODELS FOR DATA-EFFICIENT SEQUENTIAL DECISION MAKING

Nikhil Malhotra

原始摘要(英文原文)· Original abstract
Deep reinforcement learning (DRL) can achieve strong sequential decision-making performance, yet many methods remain interaction-intensive and can adapt slowly when environment dynamics, objectives, or task semantics change.We propose Dynamic Symbolic Deep Reinforcement Learning (DSDRL), a neuro-symbolic framework that couples a persistent symbolic state representation with an adaptive world model, model-based trajectory imagination, and state-differentiated reward shaping.The central hypothesis is that explicitly representing task-relevant changes between consecutive symbolic states can improve sample efficiency and post-change adaptation relative to policies trained only from externally specified scalar rewards.In contrast to symbolic reinforcement-learning architectures in which symbolic planning primarily schedules learned subtasks, DSDRL uses symbolic state as a shared interface across perception, dynamics learning, planning, reward computation, and control.The world model is updated when observed transitions disagree with predicted transitions, allowing the agent to revise its internal dynamics as interaction reveals new Nikhil Malhotra https://iaeme.com/Home/journal/IJAIML
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

DYNAMIC SYMBOLIC DEEP REINFORCEMENT LEARNING STATE-DIFFERENTIATED REWARDS AND ADAPTIVE WORLD MODELS FOR DATA-EFFICIENT SEQUENTIAL DECISION MAKING — 科研速览 Science Skim