Krzysztof B. Beć, Justyna Grabska, Vanessa Moll, Christian W. Huck
Surface scattering is a major confound in near-infrared (NIR) analysis of particulate foods. In roasted coffee powders, inhomogeneous scattering obscures cultivar differences. Standard practice eliminates scattering (extended/multiplicative scatter correction) via reference-anchored polynomial projection; in heterogeneous matrices this is order-sensitive and can remove analyte-relevant variance. We instead encode scattering explicitly; per-spectrum polynomial slope, curvature and cubic baseline coefficients are determined and appended as descriptors and models are trained on the augmented matrix. Using 300 diffuse-reflectance FT-NIR spectra (10 000–4 000 cm -1 ) from 25 lots covering 7 Arabica cultivars, this strategy improved test-set authentication; Linear Discriminant Analysis (LDA) increased from 70.8% to 83.3%; Support Vector Machine (SVM)-linear from 63.9% to 84.7%, while Quadratic Discriminant Analysis (QDA) remained high (88.9%). The approach provides a quantitative scattering-aware method, treating it as information rather than aiming for its blind elimination. This method is an interpretable, easily implemented alternative and is applicable to other particulate foods (e.g., cocoa, tea, spices). • Residual scattering curvature dominates NIR spectra of roasted Coffea arabica powders. • Conventional MSC/EMSC (implicit removal) is unstable in inhomogeneous food matrices. • Polynomial coefficients used as explicit scatter descriptors improve classification accuracy. • Scatter-descriptor models yield error patterns consistent with known cultivar relationships. • The approach is conceptually implementable in NIR analysis of other particulate foods.