Threat-Intel RAG

Ask how to detect, mitigate or recognise an attack technique — in English or Chinese — and get an answer where every sentence cites the MITRE ATT&CK passage it came from, or an explicit refusal when the data can't support one. Built to measure whether the answers can be trusted.

Recorded evaluation results — not a live model

How a question is answered

Red boxes are the points where the system refuses instead of answering.

1 · Understandrewrite revoked IDs, find named IDs, detect intent
2 · RetrieveBM25 + multilingual dense (bge-m3), 50 each
3 · MergeReciprocal Rank Fusion, named technique first → top 5
Gate 1 · Relevancebest cosine below 0.44 → refuse, no LLM call
4 · AnswerLLM returns JSON claims, each citing passage IDs
Gate 2 · Answerable?the model says the passages don't answer it
5 · Check citationsdrop claims citing passages that weren't retrieved

Explore recorded examples

Each example is a real question from the evaluation set, with the answer the model gave in the recorded run, the faithfulness judge's verdict on every sentence, and the retrieval trace. Click a citation or a row to read the passage.

What the evaluation shows

Findings
  • Retrieval is the bottleneck. Answers cited the correct passage exactly as often as retrieval found it, in every category.
  • Structured output is the main injection defense. Injected instructions have no field to go in; every claim must cite a retrieved passage.
  • Poisoning is the open risk. A poisoned passage is false data, not an instruction; with defenses on, the fake product never reached an answer, but answers still cited the poisoned copy.
Limitations
  • Evaluation sets are small and hand-written by the author (78 questions, 70 attacks per configuration).
  • Faithfulness is scored by an LLM judge, checked against 40 human labels.
  • One model (Qwen); results depend on the host serving it.
  • Questions about which groups or software use a technique are refused by design.