科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Radiological physics and technology2026-09-03

RadTechBERT: a domain-specific language model for radiological technology evaluated using Japanese national examination-derived cloze questions.

Ayako Yagahara, Noriya Yokohama, Masahito Uesugi, Mitsuhiro Aizawa, Daisuke Ando, Tomoki Ishikawa, Masaru Sudo, Naoki Nishimoto, Takumi Tanikawa, Yousuke Aoki, Yuji Tani, Takaaki Banno, Minoru Kawamata

原始摘要(英文原文)· Original abstract
Specialized language models can improve performance in expert domains; however, in radiological technology, evaluation datasets and models tailored to radiological technologists' practice remain limited and are not widely available. In this study, RadTechBERT was developed by pretraining a Bidirectional Encoder Representations from Transformers on a radiological technology-specific corpus from the initial pretraining stage, without relying on continued pretraining of existing general models. To enable systematic evaluation, a new cloze question dataset was constructed from items derived from the Japanese national examination for radiological technologists, and subject labels were added to support subject-level analyses. To develop the model, Unigram and Byte Pair Encoding (BPE) tokenizers were compared with vocabulary sizes of 32 K, 50 K, and 100 K. RadTechBERT was trained under two settings: using only the domain-specific corpus and using a mixed corpus that additionally included Wikipedia. For benchmarking, baseline models pretrained on general corpora such as Wikipedia, as well as existing and medical textbook-based models, were also evaluated. Performance was assessed using Top-5 accuracy on the cloze task, both overall and by subject. RadTechBERT with BPE_32K outperformed baselines in many subjects, with more than a twofold improvement in radiation safety management and radiation measurement relative to the strongest baseline. In contrast, gains were smaller in subjects having substantial overlap with general medicine, and Wikipedia mixing did not yield consistent improvements. The optimal tokenizer and vocabulary size were subject-dependent.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

RadTechBERT: a domain-specific language model for radiological technology evaluated using Japanese national examination-derived cloze questions. — 科研速览 Science Skim