Aristeidis Karras, Leonidas Theodorakopoulos, Christos Karras, George A. Krimpas, Αναστάσιος Γιάνναρος, Charalampos-Panagiotis Bakalis
In this work, we present a principled framework for the deployment of Large Language Models (LLMs) in enterprise big data management across digital governance, marketing, and accounting domains. Unlike conventional predictive applications, our approach integrates LLMs as auditable, sector-adaptive components that robustly and directly enhance data curation, lineage, and regulatory compliance. The study contributes (i) a systematic evaluation of seven LLM-enabled functions—including schema mapping, entity resolution, and document extraction—that directly improve data quality and operational governance; (ii) a distributed architecture that deploys Apache Spark orchestration with Markov Chain Monte Carlo sampling to achieve quantifiable uncertainty and reproducible audit trails; and (iii) a cross-sector analysis demonstrating robust semantic accuracy, compliance management, and explainable outputs suited to diverse assurance requirements. Empirical evaluations reveal that the proposed architecture persistently attains elevated mapping precision, resilient multimodal feature extraction, and consistent human supervision. These characteristics collectively reinforce the integrity, accountability, and transparency of information ecosystems, particularly within compliance-driven organizational settings.