Yujiao Wang, Chenyi Cai, Xiao Wang, Xinhua Zhu, Peng Tang
Fueled by data-intensive science, urban morphology studies have shifted from small-scale case studies to large-scale pattern recognition, thus an effective, machine-readable representation of complex urban forms is therefore critical. Beside classic metrics-based representation with morphological indicators, vision-based representation extracted by deep learning from images have gained prominence, but their relative strengths remain unexplored. This study implements a unified pipeline that constructs both representations across four morphological perspectives—street, building, landscape, and context—and evaluates their efficacy through clustering based on K-Means and Self-Organizing Map (SOM), supplemented by expert-validated case retrieval. Clustering reveals complementary logics: morphometrics segment cases along quantitative thresholds yet often group visually dissimilar forms, while vision-based models cluster perceptual patterns but sacrifice numeric precision. Expert retrieval indicates a modest but consistent tilt: vision-based result holds a slight edge in street-network and building queries, whereas metric-based representations led landscape and context queries. This study also demonstrates the efficacy of vision-based representations compared with metric-based ways and their operational trade-offs (interpretability, compute, data demand). The comparative framework guides researchers in selecting representation strategies tailored to various urban-form perspectives, enhancing the accuracy and applicability of large-scale morphological analyses for planning and design.