Sophie Robert-Hayek, Andrew Hardwick, Davide D’Amico, Frédérique Rey
Textual criticism, the task of reconstructing a text as close as possible to its original form based on multiple ancient copies, is a central concern of philologists studying ancient literature.After identifying and cataloging the available manuscripts, scholars align the texts to produce a list of variants, a process known as collation.Automating this task has been one of the earliest areas of interest for applying computational methods to the humanities.Most current approaches rely on sequence alignment algorithms, originally developed for bioinformatics.However, recent advances in mono and multilingual neural alignment, capable of modeling semantic, syntactic, and grammatical relationships, remain largely underutilized in this field.In this paper, we introduce ALMA (Alignment and Learning for Manuscript Analysis).This new pipeline integrates lexical, grammatical, and semantic analysis to perform automatic textual alignment.We evaluate ALMA against three established alignment methods, each representing a different paradigm, on four manuscripts of the Gospel of John, two in Latin and two in Greek.ALMA achieves mean alignment scores of 94.12% (Latin) and 95.77% (Greek), significantly outperforming all baseline methods: alignment scores are improved over a deep learning-only approach by 1.93% and 1.09%, over a grammar-based approach by 23.39% and 10.14%, and over a sequence alignment method by 38.57% and 13.39%, respectively, for Latin and Greek.Notably, in this study, ALMA is the only approach to achieve a perfect median of 100% success on both datasets.These results highlight ALMA's substantial potential for advancing computational approaches for textual collation and more generally for computational philology.