科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ AI-Linguistica Linguistic Studies on AI-Generated Texts and Discourses2026-05-17· Computer science

Automating Semantic Annotation in Low-Resource Languages: Evaluating GPT-4 for Urdu NLP

Gohar Rahman

原始摘要(英文原文)· Original abstract
Semantic annotation is a fundamental yet labor-intensive process essential for building effective Natural Language Processing (NLP) systems, particularly for low-resource languages such as Urdu. The limited availability of large, manually annotated datasets has constrained advancements in Urdu NLP. This study explores the potential of automating semantic annotation using GPT-4, a state-of-the-art large language model (LLM), through structured prompt engineering without task-specific fine-tuning. A corpus of 50,000 Urdu sentences spanning news articles, social media posts, and literary texts was used to evaluate three core tasks: Named Entity Recognition (NER), semantic similarity, and sentiment analysis. GPT-4 demonstrated strong performance, achieving an F1-score of 92% for NER, a Pearson correlation of 0.87 for semantic similarity, and an accuracy of 88% with a macro-F1 of 87% for sentiment classification. These results indicate that LLMs guided by instruction-based prompts can reliably perform complex NLP tasks in low-resource contexts. Nonetheless, challenges with idiomatic expressions, sarcasm, and rare entities highlight the need for carefully designed prompts and potential human-AI collaboration.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Automating Semantic Annotation in Low-Resource Languages: Evaluating GPT-4 for Urdu NLP — 科研速览 Science Skim