Quoc-Viet-Anh Tran, Gia-Bao Duong, Huu-Huy-Hoang Tran, Thi-Hai-Yen Vuong, Hoang-Quynh Le
Substance use information in clinical narratives is clinically important but is typically documented in unstructured and informal text, making automatic extraction challenging, particularly in low-resource and non-English settings. The BioCreative IX ToxHabits shared task addresses this problem in Spanish clinical text by identifying substance use triggers and their arguments. In this paper, we introduce I-NOWJ, which addresses these challenges through two complementary strategies: a low-resource strategy based on domain knowledge expansion, controlled data augmentation with quality-aware self-training, and an adaptive context selection mechanism that determines the appropriate contextual scope for each instance. Experimental results on the ToxHabits benchmark show that I-NOWJ achieves an F1 score of 94.93% on trigger detection and 91.30% on argument extraction. These results demonstrate the effectiveness of knowledge-guided modelling and adaptive context selection for robust clinical named entity recognition in low-resource settings. Our implementation and resources are publicly available at: https://github.com/candleMind/toxhabit.