Erkan Caner Ozkat
The Antioxidant Food Table is the largest open collection of measured food antioxidant capacity. It covers 3139 products assayed by the ferric reducing ability of plasma (FRAP) method. The table has served mainly as a dietary lookup source and has never been machine-readable or benchmarked. Here it was extracted into a validated open dataset (3135 records, 99.9%). A leakage-aware benchmark of antioxidant capacity prediction from product description and category was then constructed. Eighteen predictors, from naïve baselines to deep networks and fusions, were evaluated under two partitioning regimes with the same five seeds and permutation controls. Since 39.4% of records share a product name, the conventional random split rewards memorization. A learning-free duplicate lookup explained 45% of the apparent k-nearest-neighbor advantage over the category median. In grouped evaluation, ridge regression on term frequency-inverse document frequency (TF-IDF) features (R2=0.674) outperformed both deep networks. Pretrained word vectors did not close this gap. An equal-weight fusion of all eight models performed best (R2=0.684) with 2.6-fold lower variability. Protocol and representation, rather than architecture, dominated the outcome on this benchmark. The dataset, code, and predictions are released openly.