Richard G. Brereton
In the previous article, we discussed the enormous increase in the impact of chemometrics methods over the last four decades [1] and the important role PLS (partial least squares or projection to latent structures) has had in this revolution. However, we are yet to describe this technique, which will be the subject of this and subsequent articles. There are many thousands, or perhaps tens of thousands, of theoretical, methodological and tutorial articles about PLS over the last 50 years, and possibly many hundreds of thousands of articles involving the use of this approach. In the early decades of the development of chemometrics as a coherent discipline in the 1980s and 1990s, there was a significant focus on PLS, but still after so many decades, it still spawns new insights. There are conferences dedicated to PLS. This article is therefore only one of very many such articles, but PLS can be approached in endless ways, and no general introduction to chemometrics is complete without describing this method. PLS can be approached in endless ways and no general introduction to chemometrics is complete without discussing this method. PLS was first proposed in the 1960s by Herman Wold [2, 3]. The method was slowly introduced to chemometrics with a significant expansion in interest in the 1980s. Svante Wold first publicised its applicability in the 1970s and 1980s [4, 5]. Early pioneers of the 1980s include Paul Geladi [6], Harald Martens and Tormod Naes [7] who wrote classical articles/books that to this day are still viewed as essential reading. During the 1980s, there were numerous conferences, software developments and courses on PLS. This development was not only important in chemistry but also in economics and social sciences. The original PLS algorithm, called PLS1, was enhanced during this period, most notably by PLS2 but also by many other developments, which continue to this day. New theoretical articles on the properties of PLS continue as topical areas for research. As originally described, PLS was used for quantitative regression or calibration, sometimes distinguished by the terminology PLSR (PLS regression), where the c block was a continuous variable, such as a concentration, reaction rate or activity. Most of the early applications in chemistry were, for example, in NIR spectroscopy, where the aim was to calibrate the spectra to the concentration of an analyte or a class of compounds on a continuous scale. However, over the past few years, PLSDA (PLS discriminant analysis) [12] has become an important technique used for multivariate classification. In this case, the c block is discrete, representing a numerical label or a classifier. Typically, if there are two groups, c = +1 for group A, and c = −1 for group B. For multiple groups, there are several modifications [13] available. In subsequent articles, we will look at the properties of the matrices obtained using the PLS1 algorithm and how they fundamentally differ from PCA, even though some have the same names. The author declares no conflicts of interest. Data sharing is not applicable to this article as no datasets were generated or analysed during the current study.