科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in artificial intelligence2026-01-01

The legalization of international instruments: a hybrid RAG scoring framework based on chain-of-thought prompting.

Yan Chen, Zihua Zeng, Muhamad Sayuti Hassan

一句话结论 · In one sentence

GPT-5.2 and GPT-4o, based on this architecture, significantly outperform the baseline prompting strategy, traditional machine learning methods (TF-IDF+LR), and the fine-tuned pre-trained language model (Legal-BERT) on the international instrument scoring task. Among all evaluated models, GPT-5.2 exhibits the strongest overall performance. Based on the averaged outcomes of three independent runs, the model attained QWK scores of 0.806 and 0.788, and MAE scores of 0.089 and 0.084, on the ASEAN instrument test set and the test set of other major regional organizations, respectively, demonstrating a substantially high degree of agreement with human expert ratings.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Accurately evaluating the degree of legalization of international instruments is analytically valuable for assessing the level of institutionalization of interstate cooperative arrangements. However, conventional assessment approaches rely heavily on specialized legal expertise, and the inherent efficiency constraints of manual analysis make systematic evaluation across large-scale instrument corpora exceedingly difficult. METHODS: To address the lack of automated scoring tools in this domain, this study introduces a hybrid retrieval-augmented generation (RAG) scoring framework based on chain-of-thought (CoT) prompting. Based on the legalization conceptual framework proposed by Abbott and Snidal, this study constructs a five-level coding standard that covers the three dimensions of obligation, precision, and delegation, and accordingly designs a stepwise binary decision process to develop the corresponding CoT prompt template. In parallel, this study constructs a vector-indexed clause database comprising 2,611 expert-annotated samples from the ASEAN instrument corpus, providing retrieval-augmented semantic references for the scoring task. The retrieval mechanism incorporates an adaptive quality-weighted ranking strategy and employs an additional pre-trained model to perform secondary filtering of candidate results. To evaluate the effectiveness of the scoring framework, two independent test sets were constructed from ASEAN instruments and documents from other major regional organizations, respectively, containing 254 and 255 expert-annotated samples. RESULTS: GPT-5.2 and GPT-4o, based on this architecture, significantly outperform the baseline prompting strategy, traditional machine learning methods (TF-IDF+LR), and the fine-tuned pre-trained language model (Legal-BERT) on the international instrument scoring task. Among all evaluated models, GPT-5.2 exhibits the strongest overall performance. Based on the averaged outcomes of three independent runs, the model attained QWK scores of 0.806 and 0.788, and MAE scores of 0.089 and 0.084, on the ASEAN instrument test set and the test set of other major regional organizations, respectively, demonstrating a substantially high degree of agreement with human expert ratings. DISCUSSION: These findings indicate that integrating CoT prompting with the RAG pipeline improves LLM performance in scoring the degree of legalization of international legal documents. This provides methodological and technical foundations for large-scale institutional empirical analysis in ASEAN and broader legal systems, and points toward future extensions of the framework to other regional and multilateral legal orders.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The legalization of international instruments: a hybrid RAG scoring framework based on chain-of-thought prompting. — 科研速览 Science Skim