科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ ACM Transactions on Multimedia Computing Communications and Applications2026-01-05· Closed captioning

TextCoT: Zoom-In for Enhanced Multimodal Text-Rich Image Understanding

Bozhi Luan, Hao Feng, Hong Chen, Yonghui Wang, Wengang Zhou, HouQiang Li

原始摘要(英文原文)· Original abstract
The advent of Large Multimodal Models (LMMs) has fueled extensive research due to their sophisticated reasoning capabilities. However, for understanding text-rich images, challenges persist in fully leveraging the potential of LMMs, and existing methods struggle with effectively processing high-resolution images. Addressing this, we introduce TextCoT, a training-free Chain-of-Thought framework that improves text-rich image understanding by leveraging LMMs’ captioning abilities for global context and detailed local textual analysis. TextCoT comprises three stages: Global Context Generation, Macro-Scale Positioning, and Fine-Grained Visual Inspection—each contributing to a comprehensive understanding and precise information extraction needed for accurate question-answering. Our method requires no additional training, offering immediate plug-and-play functionality. We have demonstrated TextCoT’s effectiveness and adaptability across various benchmarks. The source code is available at https://github.com/bzluan/TextCoT .
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

TextCoT: Zoom-In for Enhanced Multimodal Text-Rich Image Understanding — 科研速览 Science Skim