He Huang, Yingli Gao, Bin Tian, Zhuo Yang
Data-driven prediction of crumb-rubber modified asphalt (CRMA) binder properties is usually reported on datasets in which a small number of mix-designs is swept across several test temperatures, so that the row count greatly exceeds the number of independent experiments. This study asks what such a dataset can actually support. Using 221 laboratory measurements drawn from 20 independent CRMA mix designs, we evaluate a task-adaptive mixture-of-experts multi-task network (TA-MoE-MTL), a locked Huber-anchored hybrid extension, and fourteen deep and classical reference models for the joint prediction of penetration, softening point, ductility and rutting factor. Three methodological elements are introduced. First, model selection is made strictly nested and group-aware: the stopping epoch is chosen on an inner split of the training designs and the held-out designs are used once. Second, we bound what the recorded inputs can explain before any model is fitted: because the consistency targets are constant within a mix design and because eight designs share identical input vectors while their rutting factors differ, the attainable coefficient of determination for the rutting factor is 0.790 rather than unity. Third, performance is reported with each mix design weighted equally, so that high replication designs cannot dominate. The locked hybrid assigns 90% weight to a Huber-anchored robust expert branch and 10% to a freshly trained TA-MoE-MTL branch. It attains pooled out-of-fold coefficients of determination of 0.613/0.594/0.878/0.685 and ranks first of 16 models at a mean pooled R2 0.693, exceeding MLP-sklearn (0.656) by 0.037. Across three seeds, seeds 42/43/44 give mean pooled R2 values of 0.693/0.689/0.693 (mean 0.691 ± 0.002), and all three runs remain above the frozen MLP sklearn reference. A learning curve over the number of training designs is still rising at the largest size the data allow. The contribution of this work is an evaluation protocol for replicated mix design datasets, a way of bounding their information content, and a robust hybrid that exposes rather than hides the value of a simple small-sample anchor.