Sara-Päivi Paukkeri, Tapio Frantti
Large language models (LLMs) have recently been utilized in several cybersecurity-related tasks. They have the potential for significant impact, both good and evil. Despite the widespread interest, a basic understanding of LLMs in the network traffic analysis of Internet of Things (IoT) is still lacking. This study analyzed the strengths and weaknesses of Llama 3.1 in detecting various cyberattack techniques from IoT network traffic. We focused on identifying the reasons behind the successful or unsuccessful answers of the LLM. Moreover, we compared five local LLMs to examine whether the main findings from the Llama 3.1 analysis apply to other models. This knowledge is essential for planning future research and real-life solutions, ensuring that the quality of LLM responses can be improved. Overall, LLMs demonstrate general, albeit limited, knowledge of IoT network traffic context. They are aware of various cyberattack techniques and can detect various signs of attacks in network traffic data. LLMs perform well in identifying language-based indicators of malicious activity, such as SQL queries. However, local LLMs have not demonstrated reliability in detecting cyberattacks. Moreover, they cannot process hexadecimal-formatted packets, and their success rate varies across attack types. Encrypted traffic data also weakens model performance by causing hallucinations.