Sheng Hong, Xuanqi Wang, Zeyu Mei, Thisura Bojitha Wickramaratne
With the rapid development of multimodal large language models (MLLMs), structured event extraction (EE) has emerged as a critical intelligent information processing task, with increasing demand across multilingual and multimodal application scenarios. However, significant challenges remain in zero-shot multimodal and cross-language scenarios, including inconsistent cross-language outputs and the high computational cost of full-parameter fine-tuning. This study takes VideoLLaMA2 (VL2) and its improved version VL2.1 as the core models, and builds a multimodal annotated dataset covering English, Chinese, Spanish, and Russian (including 5,728 EE samples). It systematically evaluates the performance differences of zero-shot learning, and parameter-efficient fine-tuning (QLoRA) techniques. The experimental results show that for EE, QLoRA fine-tuning yields substantial gains across both models: VL2.1 achieves the highest trigger accuracy at 65.48%, while VL2 achieves the highest argument accuracy at 60.54%, with each figure representing the best performance obtained across the two fine-tuned models respectively. The study confirms that fine-tuning significantly enhances model robustness.