科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Reliability2026-01-01· Computer science

COSTAR: Software Code Smell Detection Through Tree-Based Abstract Representation

Praveen Singh Thakur, Mahipal Jadeja, Satyendra Singh Chouhan, Santosh Singh Rathore

原始摘要(英文原文)· Original abstract
Code smells are suboptimal code structures that increase software maintenance costs and are challenging to detect manually. Researchers have explored automatic code smell detection using Machine Learning (ML) methods, which rely heavily on static code metrics or source code representation. Static code metrics often rely on structural attributes such as lines of code, cyclomatic complexity, or comment density. However, these metrics do not always reflect true code complexity and provide only quantitative insights without inherently detecting poor coding practices. In contrast, representations like Abstract Syntax Trees (ASTs) focus on the structural and syntactic elements of code, capturing hierarchical and contextual relationships within the source code. This enables precise identification of code structures such as loops, function calls, and conditionals, which are essential for detecting code smells. This paper introduces COSTAR (Code Smell Detection through Tree-based Abstract Representation), a source code representation technique using Abstract Syntax Trees (AST) to uniquely represent each source code instance. COSTAR captures the hierarchical structure of the source code by extracting all paths from the root to individual nodes within the AST. By employing a pretrained Sentence-BERT (SBERT) embedding model, COSTAR generates vectors for each extracted path. The subsequent calculation of the mean of these vectors yields a precise and comprehensive source code representation. Extensive experiments were conducted to validate COSTAR's performance using various ML techniques on four benchmark MLCQ code smell datasets: Data Class, God Class (Blob), Feature Envy, and Long Method. Various performance metrics have been employed to evaluate the model's performance. The experimental results indicate that COSTAR enhances the performance of the code smell detection model compared to existing methods. An improvement in the f1-score ranging from 0.03 (Long Method) to 0.19 (Feature Envy) was observed. Furthermore, a comparison of COSTAR with state-of-the-art methods demonstrated that it outperformed approaches like Code2Vec and CuBERT in code smell detection.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

COSTAR: Software Code Smell Detection Through Tree-Based Abstract Representation — 科研速览 Science Skim