Ali Riza Durmaz, James D. Lamb, McLean P. Echlin, Tresa M. Pollock
This work comprehensively assesses the capabilities of recently emerging foundation models for multimodal image matching and registration in the materials science and engineering (MSE) domain. A distinctive feature in correlative materials microscopy is that the images, which need to be spatially associated, commonly not only have limited mutual information but also distinct length scales. To date, it is largely unknown how foundational multimodal matchers, such as MatchAnything RoMa or ELoFTR, which were trained predominantly with macroscale images from the human environment, generalize to image pairs in such an out-of-domain setting—especially given the pronounced information and field of view disparities in image pairs. To evaluate these models, we use the recently published AmalgaMatch dataset which covers 187 image pairs partitioned into six groups which represent distinct image matching tasks commonly faced in correlative materials microscopy and 19 subsets representing different materials. This dataset with its diverse microscopy modalities and broad range of depicted metals, alloys, and ceramics facilitates evaluation of these models in a rather representative manner. We observe that MatchAnything RoMa, as opposed to the ELoFTR variant, attains satisfactory matching results for many matching tasks and materials represented in this dataset. In contrast, image pairs within the dislocation characterization, slip partitioning and multi-scale groups pose a difficult challenge for these models. These pairs often exhibit limited mutual information and strong field of view mismatch, with field of view ratios reaching down to 2%. We propose inference strategies and workflows which embed a foundation model in an image processing pipeline to increase the matching quality and robustness and ultimately overcome aforementioned challenges inherent to MSE matching tasks.