Engineering and Technology Horizons | 2026
Authors: Dao-Duy T.
DOI: 10.55003/ETH.430206
Journal: Engineering and Technology Horizons
Year: 2026
Publisher: King Mongkut's Institute of Technology Ladkrabang
Document Type: Article
Open Access: All Open Access; Gold Open Access
Cited by: 0
Large Language Models (LLMs) are increasingly utilized for automated Cyber Threat Intelligence (CTI) tasks, such as vulnerability analysis and security advisory generation. However, LLMs are susceptible to hallucination, which refers to the generation of plausible yet factually incorrect content, posing significant risks in security-critical contexts. Although concerns have increased, there is currently no dedicated benchmark for the systematic evaluation of hallucination in LLM-generated cyber threat intelligence (CTI). This study introduces HalluCVE, a multi-signal benchmark designed to detect hallucinations in LLM-generated Common Vulnerabilities and Exposures (CVE). HalluCVE incorporates four complementary detection components: 1) Natural Language Inference-based entailment scoring, 2) lexical factual alignment, 3) LLM-as-a-Judge self-reflection, and 4) cross-model consensus divergence. Five state-of-the-art LLMs are evaluated on 1000 CVE entries as the dataset, from 2022 to 2026, encompassing both known (pre-training cutoff) and unknown (post-cutoff) vulnerabilities. The results indicate pervasive hallucination across all models, with mean Hallucination Index values ranging from 0.480 to 0.820. Notably, models demonstrate near-universal confabulation, reaching up to 100%, when queried about post-cutoff vulnerabilities, and frequently respond with high confidence instead of appropriate refusal. HalluCVE establishes a rigorous evaluation framework for assessing LLM reliability in security-sensitive CTI applications and provides insights into potential mitigation strategies. © 2026, King Mongkut's Institute of Technology Ladkrabang. All rights reserved.
Benchmarking; Cyber Threat Intelligence; Hallucination Detection; Large Language Models