What was tested
Nakato’s headline figures come from an internal set of 500,000 financial-services test cases, and every comparator for them was run on the same cases, so those comparisons are like for like; the figures are cleared for publication. The reasoning results further down come from two established benchmarks, DROP and HotpotQA, and a comparison is drawn only where the same benchmark produced one. Results vary with the kind of document and how Nakato is configured.