Christian Goldschmied
2026
Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges
Daniel Schlör | Marius Bohn | Maximilian Wolf | Kevin Bergner | Christian Goldschmied | Andreas Hotho
Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26
Daniel Schlör | Marius Bohn | Maximilian Wolf | Kevin Bergner | Christian Goldschmied | Andreas Hotho
Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26
Evaluating Large Language Model (LLM) agents on complex multi-step cybersecurity tasks requires structured, reproducible evaluation rubrics. We present BraceGreen, a framework that formalizes Capture-the-Flag (CTF) attack paths as knowledge graphs and uses them as gold-standard rubrics for agentic LLM evaluation. Each node in our knowledge graphs represents an attack step annotated with MITRE ATT&CK tactics, goals, commands, expected outputs, and semantic outcomes, while edges encode prerequisites, dependencies, and alternative paths. Our LangGraph-based evaluation workflow employs LLM-as-judge with chain-of-thought reasoning to semantically compare agent predictions against knowledge graph-encoded alternatives. We contribute a benchmark of 7 CTF machines with knowledge graph annotations, three evaluation modes (command prediction, goal inference, anticipated result), and integration with live machine infrastructure via virtual machines and a MCP server. Our approach bridges the gap between unstructured CTF writeups and graph-structured evaluation rubrics.
2025
TransAlign: Machine Translation Encoders are Strong Word Aligners, Too
Benedikt Ebing | Christian Goldschmied | Goran Glavaš
Findings of the Association for Computational Linguistics: EMNLP 2025
Benedikt Ebing | Christian Goldschmied | Goran Glavaš
Findings of the Association for Computational Linguistics: EMNLP 2025
In the absence of sizable training data for most world languages and NLP tasks, translation-based strategies such as translate-test—evaluating on noisy source language data translated from the target language—and translate-train—training on noisy target language data translated from the source language—have been established as competitive approaches for cross-lingual transfer (XLT). For token classification tasks, these strategies require label projection: mapping the labels from each token in the original sentence to its counterpart(s) in the translation. To this end, it is common to leverage multilingual word aligners (WAs) derived from encoder language models such as mBERT or LaBSE. Despite obvious associations between machine translation (MT) and WA, research on extracting alignments with MT models is largely limited to exploiting cross-attention in encoder-decoder architectures, yielding poor WA results. In this work, in contrast, we propose TransAlign, a novel word aligner that utilizes the encoder of a massively multilingual MT model. We show that TransAlign not only achieves strong WA performance but substantially outperforms popular WA and state-of-the-art non-WA-based label projection methods in MT-based XLT for token classification.