Mounir Basbas, Badis Djamaa, Mustapha Réda Senouci
ABSTRACT The rising rate of cyberattacks and the dearth of skilled cybersecurity professionals require innovative solutions beyond traditional security measures. Large Language Models (LLMs), known for their natural language processing abilities, present a viable way to improve cybersecurity defenses. However, despite LLMs' exploration in various security applications, systematic research aligning LLMs' contributions with established cybersecurity frameworks is lacking. To bridge this gap, this paper presents a systematic literature review aligned with the NIST Cybersecurity Framework (CSF 2.0) to gain a clear understanding of LLMs' multifaceted contributions and uncover their overlooked potential, particularly in areas such as the Awareness and Training category of the NIST CSF 2.0 Protect function. Leveraging the accessibility and privacy benefits of open‐source LLMs (OSLLMs), we developed a benchmarking methodology using Multiple‐Choice Question Answering (MCQA) to evaluate 21 OSLLMs recognized for their state‐of‐the‐art performance (e.g., Llama‐3‐70B and Qwen2‐72B, over 80% in MMLU) across two publicly available datasets. Our findings reveal that while larger LLMs generally outperformed smaller ones, medium‐sized LLMs (7B‐27B) achieved accuracy within 5–9 percentage points of 70B+ models, contesting the claim that model size alone dictates efficacy. These insights motivated a further investigation into optimized prompt engineering workflows using the DSPy framework. Our results reveal that advanced prompting techniques significantly boost the performance of smaller models, narrowing the performance gap with larger ones and broadening deployment possibilities. Furthermore, we introduced structural modifications and novel exit options to address position and forced‐choice biases in standard MCQA datasets, further enhancing the reliability of OSLLM evaluations. Overall, the results of this study not only strengthen our understanding of OSLLMs' aptitude in cybersecurity but also emphasize the impact of advanced prompt engineering and unbiased datasets in enhancing the capabilities of OSLLMs in cybersecurity.