Xinkang Chen, Zhou He
Handwritten mathematical expression recognition (HMER) aims to convert images of handwritten mathematical expressions into structured LaTeX sequences. Although recent encoder-decoder models have achieved strong performance, they still struggle with visually similar symbols, Greek letters, special symbols, and structural tokens such as fractions, radicals, superscripts, and subscripts, especially when expressions contain complex spatial layouts. To address these issues, we propose a Counting-based Position-aware Transformer (CP-Former) that integrates symbol-counting supervision and position-aware structural modeling into a CoMER-based recognition framework. Instead of treating counting and position information as isolated cues, CP-Former jointly learns global symbol occurrence statistics and relative structural positions from LaTeX annotations, without requiring additional symbol-level labels. Experiments on CROHME 2014, 2016, and 2019 show that CP-Former consistently improves over the CoMER baseline, with absolute ExpRate gains of 3.00%, 0.96%, and 2.00%, respectively. Compared with recent state-of-the-art methods such as PosFormer, CP-Former achieves competitive exact-match performance and obtains comparable or better results on several error-tolerant metrics, particularly under the ≤2 and ≤3 settings. These results suggest that the proposed counting-based position-aware modeling improves structural robustness for complex handwritten expressions, whereas the fair ablation study further shows that the gains cannot be attributed to counting supervision alone, but mainly arise from its integration with position-aware structural modeling.