Zekun Wang
"This dataset provides 19,468 question-answer-evidence metadata records for evaluating verifiable text-rich image understanding under visual degradation. The records are derived from the test splits of Total-Text (8,065 records) and SCUT-CTW1500 (11,403 records) and cover region reading, spatial relations, attribute retrieval, counting, and evidence selection. Each record specifies a question, reference answer, supporting text instance or instances, polygon and bounding-box geometry, difficulty tags, and task-specific construction metadata. The release contains only derived metadata; it does not redistribute source images, original annotation files, degraded views, predictions, model weights, or source code. Users must independently obtain source images under the original dataset terms. A JSON schema, field dictionary, construction criteria, and checksums are included to support reproducible loading and validation. The dataset supports evaluation of answer correctness together with grounding and evidence-text consistency."