Ayan Myssayev, Amin Tamadon, Nadiar M Mussin, Reza Shirazi, Afshin Zare
Artificial intelligence-based cytology classification achieves high diagnostic accuracy for benign-malignant discrimination and shows promising but less mature performance for histological subtyping. Prospective external validation, standardized reporting, and clinically integrated human-artificial intelligence studies are required before routine clinical deployment.
BACKGROUND: Cytology specimens are essential for lung cancer diagnosis, yet distinguishing benign from malignant cells and subtyping adenocarcinoma, squamous cell carcinoma, and small-cell lung cancer remains challenging due to limited cellular material.
METHODS: We systematically searched seven databases for studies evaluating artificial intelligence-based lung cancer classification in cytology. After screening, 32 studies were included, of which 20 contributed 33 analyzable two-by-two data rows for diagnostic test accuracy meta-analysis. Outcomes were pooled using bivariate logit random-effects models. Summary receiver operating characteristic curves, Fagan nomograms, Deeks' test for publication bias, and quality assessments using QUADAS-2 and artificial intelligence-specific tools were performed.
RESULTS: For benign versus malignant diagnosis, 18 data rows encompassing 3464 observations yielded a pooled sensitivity of 89.3% (95% confidence interval, 84.1-92.9), a specificity of 88.5% (95% confidence interval, 83.2-92.3), and an area under the curve of 0.952. At a median prevalence of 58.2%, the post-test probability of malignancy after a positive artificial intelligence result was 91.2%, whereas a negative result yielded a post-test probability of 14.6%. For one-versus-rest subtyping among adenocarcinoma, squamous cell carcinoma, and small-cell lung cancer, 14 data rows comprising 2687 observations demonstrated a pooled sensitivity of 77.5%, specificity of 89.5%, and area under the curve of 0.901. Small-cell lung cancer showed the strongest performance (specificity 93.2%, area under the curve 0.949), whereas squamous cell carcinoma had lower sensitivity (67.0%). Substantial heterogeneity was observed across studies. Deeks' test suggested potential small-study effects, and the risk of bias was high in all included studies.
CONCLUSION: Artificial intelligence-based cytology classification achieves high diagnostic accuracy for benign-malignant discrimination and shows promising but less mature performance for histological subtyping. Prospective external validation, standardized reporting, and clinically integrated human-artificial intelligence studies are required before routine clinical deployment.