M.-E. Hoeppli, S. Pascual-Diaz, E. E. Biggs, L. E. Simons, R. C. Coghill, M. Lopez-Sola, C. King, N. Aghaeepour, M. Angst, B. L. J. Gaudilliere, J. N. Stinson, M. Moayedi
Multisite functional Magnetic Resonance Imaging (fMRI) studies are rapidly becoming the norm to allow the acquisition of large samples required to perform advanced statistical techniques, e.g. machine learning. However, acquiring data at multiple sites includes methodological and technical challenges due to the difference in the facilities of each site, such as scanner manufacturer and model. These new challenges add to the already well-known challenges of fMRI data, including scanner-related noise, participant movement, and physiological confounds. To ensure the optimal quality of data, a balance needs to be achieved between standardizing data preprocessing across sites and optimizing within-site data preprocessing. To define the optimal preprocessing pipeline including standard steps and additional denoising technique for our multisite dataset, we first tested 3 commonly used preprocessing pipelines, i.e. fMRIPrep, FSL, and CONN. These pipelines all include standard steps, e.g. motion correction, temporal filter, etc. In a second step, to further improve signal quality, we tested the efficiency of 3 additional denoising techniques on the output of the data preprocessed with the previously defined pipeline. These techniques included aCompCor, FSL FIX and ICA-AROMA. Signal quality was quantified as temporal signal-to-noise ratio (tSNR) across and within sites. Our results show that the performance of the preprocessing pipeline and denoising technique varies between sites. FSL yielded the highest tSNR at two sites and fMRIPrep at the third, revealing a significant site x pipeline interaction. An FSL-based pipeline with minimal adjustment to accommodate site-specific parameters and followed by denoising using a single FSL FIX classifier, which was custom-trained across sites, achieved the highest quality of signal in our data and yielded the greatest consistency in improved data quality across sites. Because of the potential of this pipeline to adequately identify and address site-specific noise, while remaining constant across sites, we selected it as optimal preprocessing pipeline for our dataset.