科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Plant communications2026-09-08

iPheno: A Novel Dual-Aware Vision-Language Model for Multi-Task Fine-Scale Crop Phenotyping.

Laiyi Fu, Hongbo Liu, Yanbo Han, Shunkang Ling, Hongming Zhang, Fei-Yue Wang, Danyang Wu, Hequan Sun

原始摘要(英文原文)· Original abstract
Crop phenotyping is crucial for advancing plant breeding, yet remains a significant challenge. Manual approaches are labor-intensive and do not scale to the analysis of large datasets, while computational methods like Vision-Language Models (VLMs) lack the adaptability for fine-scale spatial reasoning and diverse phenotyping scenarios. To bridge the gaps, we present iPheno, a fully open-source, domain-specialized multimodal VLM for fine-scale crop phenotyping. A key innovation of iPheno is its dual-aware architecture. First, a spatial-aware feature extractor samples mask regions into local K-Nearest Neighbors (KNN) graphs and enables fine-scale analysis of arbitrary-shaped regions; second, a task-aware Mixture-of-Experts (MoE) routing mechanism activates specialized modules for each phenotyping task. To train iPheno and achieve rigorous benchmarking, we constructed iPheno-120K, a large-scale high-precision dataset designed for multiple phenotyping tasks. Evaluations on iPheno-120k test set and other publicly available datasets showed that iPheno outperformed all fine-tuned baselines, by improving F1-score by 17.1% (LLaVA-1.6-13B) to 28.9% (MiniCPM-o-9B), while achieving the highest inference speed and memory efficiency. A web server (https://ipheno.ai4bread.com), a mobile application (www.ipheno.cn), and a stand-alone PC client (https://github.com/2997029323/iPheno-PC-Client) are available for iPheno.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

iPheno: A Novel Dual-Aware Vision-Language Model for Multi-Task Fine-Scale Crop Phenotyping. — 科研速览 Science Skim