Iván Martínez-Murillo, María Miró Maestre, Armando Suárez Cueto, Elena Lloret, Paloma Moreda
Large Language Models (LLMs) have shown remarkable capabilities, demonstrating their potential to transform the Natural Language Processing (NLP) field by achieving strong performance across the entire spectrum of tasks. In this context, there is growing interest in leveraging LLMs as automated crowdworkers to streamline the traditionally labor-intensive process of manual annotation and linguistic content generation. This paper specifically examines the feasibility of using LLMs to generate high-quality contextual information — an increasingly recognized element for enhancing generative systems — by producing linguistic contexts from given premise sentences. We evaluate four prominent LLMs — LLaMA 2, PaLM 2, Vicuna, and GPT-3.5 — by making them generate 180 contextual outputs, which are then compared to 60 contexts manually crafted by linguistic experts. To systematically assess the appropriateness of these generated contexts, we introduce CATS (Contextual Appropriateness in Texts for Spanish), a novel and adaptable evaluation metric for measuring the appropriateness of contextual information generated texts. CATS is rooted in established discourse theories and provides a robust framework for analyzing linguistic context quality. The implementation of CATS is made publicly available at https://github.com/gplsi/cats . The reliability of CATS is validated through a manual evaluation conducted by two independent human referees. The results indicate that the quality of contexts generated by LLMs is comparable to those produced by human specialists, as evidenced by both CATS scores (0.214 vs. 0.150) and human judgment. Our findings underscore the efficiency and cost-effectiveness of employing LLMs as alternatives to human crowdworkers for generating contextual information, offering promising implications for the scalability and advancement of generative systems.