Wenjie Ping, Yanan Jiao, Huijie Fan, Xuguang Zhang
Financial statement fraud inflicts enormous economic damage and continues to evade conventional detection methods that rely predominantly on structured accounting ratios. Such approaches discard the rich forensic signals present in narrative disclosures and document formatting—two channels that fraudsters routinely manipulate alongside the numbers themselves. This paper introduces TriAudit, a trimodal deep learning framework that jointly encodes three complementary evidence streams from corporate 10-K filings: (i) structured financial ratios and accrual indicators processed through a residual MLP, (ii) MDA textual semantics captured via a FinBERT-based encoder with cross-chunk attention, and (iii) visual document-layout features extracted using a LayoutLMv2 backbone augmented with a novel Layout Deviation Score that quantifies spatial anomalies against industry-normative templates. These modality-specific representations are integrated through a hierarchical cross-modal attention mechanism—mirroring the auditor’s evidential reasoning from numerical signals through narrative context to document structure—followed by a gated fusion module that adaptively weights each stream based on perfiling reliability. To support auditability, TriAudit incorporates a contrastive evidence-chain module that aligns anomaly scores across modalities into structured, human-interpretable audit trails linking specific financial irregularities, linguistic deviations, and layout inconsistencies. Evaluated on over 6,000 filings from the EDGAR Corpus with fraud labels derived from SEC Accounting and Auditing Enforcement Releases, TriAudit achieves an AUC-ROC of 0.923 and F1 of 0.847, surpassing the strongest bimodal baseline by 5.5 AUC points and the classical Beneish M-score by over 18 points. Ablation experiments confirm that each architectural component—including the layout deviation scoring, hierarchical attention ordering, contrastive alignment, and orthogonality regularization—contributes meaningfully to both predictive accuracy and interpretability.