Prosenjit Sen, Toufique Ahmed, Md Sultan Mahmud, Priyanka Sarkar, Md Abidur Rahman Nibir
The article presents a comprehensive dataset of physicochemical descriptions of commercial cotton fibers (Gossypium hirsutum L.) that grow in a wide range of agro-climatic- conditions, such as Brazil, West Africa (Mali, Benin, Burkina Faso, Ivory Coast, and Cameroon), Australia, and India. It comprises paired samples from two industry-standard characterization systems, the High Volume Instrument (HVI) to measure the bundle characteristics (e.g., Micronaire, Upper Half Mean Length, and Strength) and the Advanced Fiber Information System (AFIS) to measure the distributions of single fibers (e.g., Neps per gram, Short Fiber Content, Maturity Ratio). To ensure the retention of metadata of various originating file structures, data from industrial-quality reports were digitized with a proven Python-based Optical Character Recognition (OCR) pipeline and subsequently verified manually to achieve high accuracy. The strategic importance of this dataset is that it is the aggregation of ubiquitous HVI trade values and capital-intensive AFIS process values on identical Lot and Bale identifiers. This unique synchronization offers textile-machine learning researchers or data scientists the opportunity to explore various potential linear, non-linear and interdependencies between HVI and AFIS features. Notably, it may pave the way for predicting nep count from the HVI dataset, potentially eliminating the necessity of purchasing an expensive AFIS machine. No predictive analyses or machine learning models were conducted in this current study. Besides, it can be used for various classification-based works through standard commercial classification-level data.