Ziheng Zhang, Minghui Tang, Feng Han, Hidenori Kitai, Kenji Hirata, Katsuhiko Ogasawara, Satoshi Konno, Kohsuke Kudo
Much clinically salient information in Japanese electronic medical records is recorded in SOAP (Subjective, Objective, Assessment, Plan) notes, which contain a lot of unstructured text preventing the further usage for analysis. Structuring such unstructured text is essential for secondary data use and clinical research. This study aimed to develop a schema-guided, hierarchical multi-label (15 first-level and 58 second-level labels) framework for structuring respiratory medicine SOAP notes using a locally deployed large language model (LLaMA-3.1-8B adapted with QLoRA). 200 randomly selected notes were annotated and split in two for training and test. Performance was assessed through atomic-cell matching under both strict and tolerant similarity metrics, achieved a micro-averaged F1 of 0.49 (strict) and 0.59 (Levenshtein> 0.7). Common and standardized fields such as vital signs were reliably structured, while ambiguous boundaries and rare labels remained challenging. These findings demonstrate that schema-guided hierarchical multi-label frameworks can support the structuring of Japanese clinical notes and serve as a foundation for downstream research and quality-improvement workflows.