Thomas L. Lee
Over the past few years there has been great advancements in machine learning. It is now possible to learn models that predict some quantity of interest with high accuracy. For example, we now have models which accurately predict the contents of natural images and/or the continuation of a text. However, a core part of learning is being able to adapt to newly observed data. This problem, which we will refer to as continual learning (CL), of updating a model when new information becomes accessible, is still challenging and largely unsolved. In this thesis we will approach CL from a probabilistic perspective. The advantage of doing so is that we can concretely model the data generating process, including distribution shifts. This allows for: a) the analysis of the difficulty of continually learning under different types of data generating processes and b) the design of algorithms to improve continual learning performance by leveraging a priori knowledge of the data generating process. We begin the thesis by, separating the impact of learning on a data stream with a changing distribution from the problem of only being able to access part of the stream at any point in time. We call this second problem the chunking problem, it occurs in all continual learning scenarios for which it is impossible to store all the seen data. Through experimentation, we find that, the chunking problem contributes to a significant part of the challenge of continual learning. Furthermore, we find that forgetting of previous knowledge is a key challenge in the chunking problem. This contrasts with previous work in CL, which often assume that distribution shift is the main or only cause of forgetting in CL. We also observe that, at the time, state-of-the-art CL methods tackled the distribution shift part of CL scenarios but left unaddressed the chunking problem. This suggested that a fruitful way to further improve performance in CL would be to look at the chunking problem. We took some small steps in this direction. By analysing the linear case, we found a form of model averaging to improve performance on the chunking problem. We further showed this method led to improvements in CL settings where distribution shifts are also present. After exploring chunking, we then move on to proposing methods to improve CL performance. We first start by looking at how to adapt a probabilistic classifier to shifts in the representation space it is defined on. By focusing on how to improve continual learning performance for a classifier and not the representation, which is known to forget less, we can leverage tractable Bayesian models. Specifically, we look at class-conditional Gaussian models and derive well performing update rules for the classifier and also for a memory buffer used to store previous examples. Then, we move on to looking at adapting time series foundation models. Here, we look at a reduced continual learning problem where we want to fine-tune the model to perform better on a particular type of time series. We identify that time series datasets can be seen as sampled from mixture of sub-domains. We use this insight to improve robustness to shifts in the frequency of each subdomain. We achieve this by fine-tuning individual modules on the data sampled from each sub-domain and then when forecasting using the module associated with the current most likely subdomain. We conclude the thesis by providing ideas for future work in CL.