Md Omar Faruque, Peter Jamieson, Ahmad Patooghy, Abdel-Hameed A. Badawy
Traditionally, inserting realistic Hardware Trojans (HTs) in complex hardware systems has been a time-consuming manual process, requiring comprehensive knowledge of the design and navigating intricate Hardware Description Language (HDL) codebases. Machine Learning (ML)-based approaches have attempted to automate this process but often struggle with the need for extensive training data, learning time, and limited generalizability across diverse hardware design landscapes. This paper introduces GHOST, an automated tool that leverages Large Language Models (LLMs) for rapid generation and insertion of HT. The research encompasses both the development of the GHOST framework and a comprehensive evaluation of its effectiveness across three state-of-the-art LLMs-GPT-4, Gemini-1.5-Pro, and Llama-3-70B. According to our evaluations, GPT-4 demonstrates the best performance by successfully generating and inserting HTs in 88.9% of its attempts. This study also highlights the security risks posed by LLM-generated HTs, as 100% of successful GHOST-generated HTs that completed inference within the time limit evaded detection by a state-of-the-art ML-based HT detection tool. These results underscore the need for advanced detection and prevention mechanisms in hardware security to address the emerging threat of LLM-generated HTs.