Inda Rusdia Sofiani, Hadi Suyono, Erni Yudaningtyas, Fitri Utaminingrum
Accurate acetowhite lesion segmentation is a prerequisite component of automated VIA-based screening pipelines in resource-constrained clinical settings, with potential utility in CIN boundary delineation; however, clinical translation requires validation beyond mask-based metrics.Standard transformer segmentation architectures only model attention from feature content, neglecting spatial proximity among tokens and therefore missing positional cues critical for boundary delineation.Concurrently, unified decoder designs conflate semantic and boundary channels, leaving fine boundary discrimination underaddressed.To solve these limitations, we propose SegFormer-RPB-ASPP-CSG, a parameter-efficient, boundary-aware framework that integrates three established componentsproximityaware attention, multi-scale context aggregation, and adaptive boundary-semantic decompositioninto a unified domain-specific architecture for VIA cervicography.Specifically, the Relative Position Bias Multi-Head Attention (RPB-MHA) injects learnable pairwise proximity weights into each Mix-Transformer encoder stage to enrich the attention matrix with spatial structure.In the decoder pathway, a Dilated Atrous Spatial Pyramid Pooling (Dilated ASPP) module aggregates multi-scale context, while an Upgraded Conceptual Split Gate (Upgraded CSG) prevents gradient interference between coarse lesion classification and fine contour delineation via adaptive soft gating.Evaluated on a 2,724-image VIA cervicography cohort with 681 held-out test images, the full model achieves a mean Intersection over Union (mIoU) of 87.57% and a mean Dice Similarity Coefficient (mDSC) of 90.51% using only 1.78M parameters.The per-class DSC reaches 86.66% (C0-Blue), 81.95% (C1-Lime), 93.76% (C2-Red), and 99.67% (C3-Background).Notably, a boundary-optimized RPB+CSG variant further reduces the parameter footprint to 1.47M while delivering superior boundary precision (HD95: 5.56 px vs. 5.97 px; ASD: 0.76 px vs. 0.84 px; NSD: 96.81% vs. 96.66%)and a marginal mDSC increase to 90.52%.Among the nine original comparison architectures, our framework achieves the third-highest mDSC while requiring 17.7× fewer parameters than the top performer (AUNet-MHA) and 2.53× fewer than the second-ranked method (SwinHR); among the sixteen total comparison methods (including seven retrained cervical-specific architectures), our model ranks eighth in mDSC.These results demonstrate that domainaligned integration of established architectural components can achieve competitive boundary precision within a compact parameter budget, supporting feasibility as a segmentation component in resource-constrained VIA pipelines, pending clinical validation against diagnostic endpoints.