Ask how to detect, mitigate or recognise an attack technique — in English or Chinese — and get an answer where every sentence cites the MITRE ATT&CK passage it came from, or an explicit refusal when the data can't support one. Built to measure whether the answers can be trusted.
Red boxes are the points where the system refuses instead of answering.
Each example is a real question from the evaluation set, with the answer the model gave in the recorded run, the faithfulness judge's verdict on every sentence, and the retrieval trace. Click a citation or a row to read the passage.