Baichuan Mo, Hanyong Xu, Ruoyun Ma, Jung-Hoon Cho, Dingyi Zhuang, Xiaotong Guo, Jinhua Zhao
This study evaluates large language models (LLMs) for travel behavior prediction under different levels of labeled-data availability. We compare three LLM-based frameworks: zero-shot direct prompting, textual-gradient prompt optimization from a small labeled budget, and supervised prediction using LLM text embeddings. These methods are benchmarked against multinomial logit, random forests, neural networks, and TabPFN under a budget-matched protocol on Swissmetro mode choice, London mode choice, and NHTS trip-purpose prediction. The results show a clear data-availability pattern. In scarce-label settings, direct LLM prediction is competitive with, and sometimes significantly better than, supervised/tabular baselines. Textual-gradient optimization can learn prompts that match expert hand-crafted prompts without manually encoded numerical cues, although its gains are task-dependent. As labeled budgets grow, conventional supervised and tabular models become stronger. Diagnostic tests further suggest that LLM predictions respond to supplied travel-time and travel-cost structure rather than simply memorizing benchmark records, while generated explanations should be treated as auditable but imperfect rationales.