科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in computational neuroscience2026-01-01

Explicit encoding of temporal dynamics for interpretable melody similarity modeling: insight into machine learning models.

Yilin Zhang

一句话结论 · In one sentence

Results show that the engineered temporal feature space exhibits significant nonlinearity, with the highest performance achieved by Deep GRU (R 2 = 0.863, MAE = 0.058), followed by XGBoost among non-recurrent models (R 2 = 0.824, MAE = 0.091). Baseline models showed lower performance, demonstrating the nonlinear characteristics of the task. Feature attribution analysis identified temporal descriptors and embedding-based variables as the most influential features.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Temporal melody similarity is a fundamental problem in music modelling, and current methods are mostly based on recurrent or attention-based architectures that implicitly learn sequential structures. METHODS: This work adopts an explicit feature-representation approach using temporal-style descriptors, including first-order deltas, relative ratios, polynomial interactions, and rolling statistical features. These features are represented in a model-agnostic manner to enable controlled comparison, interpretability, and reproducible evaluation across architectures. Seven model families were evaluated, including linear baselines, ensemble methods, recurrent networks, attention-based models, and dense fusion architectures, using a large-scale melody similarity dataset containing over 11,000 samples. Cross-validation and bootstrap-based uncertainty estimation were applied for performance evaluation. RESULTS: Results show that the engineered temporal feature space exhibits significant nonlinearity, with the highest performance achieved by Deep GRU (R 2 = 0.863, MAE = 0.058), followed by XGBoost among non-recurrent models (R 2 = 0.824, MAE = 0.091). Baseline models showed lower performance, demonstrating the nonlinear characteristics of the task. Feature attribution analysis identified temporal descriptors and embedding-based variables as the most influential features. DISCUSSION: The results demonstrate that TFR provides an interpretable and reproducible foundation for evaluating future sequence-learning architectures, particularly for data-constrained or deployment-oriented applications.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Explicit encoding of temporal dynamics for interpretable melody similarity modeling: insight into machine learning models. — 科研速览 Science Skim