Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges
Daniel Schlör, Marius Bohn, Maximilian Wolf, Kevin Bergner, Christian Goldschmied, Andreas Hotho
Abstract
Evaluating Large Language Model (LLM) agents on complex multi-step cybersecurity tasks requires structured, reproducible evaluation rubrics. We present BraceGreen, a framework that formalizes Capture-the-Flag (CTF) attack paths as knowledge graphs and uses them as gold-standard rubrics for agentic LLM evaluation. Each node in our knowledge graphs represents an attack step annotated with MITRE ATT&CK tactics, goals, commands, expected outputs, and semantic outcomes, while edges encode prerequisites, dependencies, and alternative paths. Our LangGraph-based evaluation workflow employs LLM-as-judge with chain-of-thought reasoning to semantically compare agent predictions against knowledge graph-encoded alternatives. We contribute a benchmark of 7 CTF machines with knowledge graph annotations, three evaluation modes (command prediction, goal inference, anticipated result), and integration with live machine infrastructure via virtual machines and a MCP server. Our approach bridges the gap between unstructured CTF writeups and graph-structured evaluation rubrics.- Anthology ID:
- 2026.kallm-1.15
- Volume:
- Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26
- Month:
- May
- Year:
- 2026
- Address:
- Palma, Mallorca (Spain)
- Editors:
- Gilles Sérasset, Katerina Gkirtzou, Michael Cochez, Jan-Christoph Kalo
- Venues:
- KaLLM | WS
- SIG:
- Publisher:
- ELRA Language Resources Association (ELRA)
- Note:
- Pages:
- 144–154
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-ws-kgllm-15
- DOI:
- 10.63317/4ne74wm45iti
- Cite (ACL):
- Daniel Schlör, Marius Bohn, Maximilian Wolf, Kevin Bergner, Christian Goldschmied, and Andreas Hotho. 2026. Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges. In Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26, pages 144–154, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
- Cite (Informal):
- Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges (Schlör et al., KaLLM 2026)