Mohammed Lataifeh, Naveed Ahmed, Imad Afyouni, Zulaiha Afrah Sadakathullah Shaduly, Abulrahman Abdulkarim
This research presents a novel design and implementation of an Intelligent Virtual Agent (IVA) in mixed reality that has advanced speech capabilities from large language models and integrates computer vision to perceive the user’s environment and actions in the real-world context. Scene understanding allows the IVA to navigate in the user’s physical space, demonstrate an understanding of the user’s actions, and dynamically interact with real-world entities. We propose a comprehensive framework for this multimodal integration, which enables the IVA to tailor its assistance and provide adaptive guidance to the users based on the actions they take in the real-world. To demonstrate and evaluate the proposed framework, we implemented two novel scenarios. Results demonstrated that participants consistently reported higher engagement, interactivity, and effectiveness with the IVA despite taking more time to complete the task. Moreover, all participants valued the IVA’s ability to adapt to their actions, offering a more personalized experience.