Measured research benchmark — not this demo
Raw retrieval
1,605
input tokens per question
62.5%
answers correct
BlockHarmony
426
92.5%
Thesis 7, our synthetic single-fact benchmark (120 questions, Claude Sonnet 4.5): average input tokens and answer accuracy per question with correctly extracted claims; with automatically extracted claims accuracy was 83.3%. In Thesis 9, our synthetic multi-hop benchmark (118 two-to-four-hop questions, same model), the same comparison was 4,823 → 591 input tokens and 71% → 97% accuracy. These are research results on specific synthetic workloads. They are not measurements of this demo and not a general guarantee.