Rodrigo Kataishi
Retrieval-Augmented Generation (RAG) systems depend on precise document retrieval to integrate external knowledge into large language models (LLMs). However, ensuring retrieval accuracy remains a challenge, particularly in corpora with overlapping topics and thematic diversity. This study introduces the concept of topic-enriched embeddings, which combine traditional term-frequency methods and advanced topic modeling techniques-such as TF-IDF, Latent Semantic Analysis (LSA), and Latent Dirichlet Allocation (LDA)-with modern SOTA contextual embeddings (all-minilm). Topic-enriched embeddings capture both term-level and topic-level semantics, using latent topic structures and dimensionality reduction to improve semantic clustering, retrieval precision, and computational efficiency. Using a legal text dataset, the proposed method demonstrates superior performance across clustering coherence and retrieval metrics. These findings underscore the potential of topic-enriched embeddings as a foundational component to improve knowledge-intensive RAG systems.