Benchmarks and method

Last updated

Nakato is measured on the answer as well as on detection: how correct the model’s answer is once each sensitive value has been swapped for a semantic twin — fictional, but faithful in everything the task depends on — and put back, and how much an active extraction attack can recover. The four headline figures are set once, on the home page. This page is what each one measures and how it was tested, with the results on reasoning beside them.

What was tested

Nakato’s headline figures come from an internal set of 500,000 financial-services test cases, and every comparator for them was run on the same cases, so those comparisons are like for like; the figures are cleared for publication. The reasoning results further down come from two established benchmarks, DROP and HotpotQA, and a comparison is drawn only where the same benchmark produced one. Results vary with the kind of document and how Nakato is configured.

Downstream utility

Downstream utility is how correct the model’s answer is after pseudonymisation and reversal, scored against its answer on the raw data. It is measured on summarisation. On the same cases, placeholder masking scores 41.3% and tokenisation 12.8%; Nakato’s figure is on the home page. Because the context survives, downstream tasks can be stacked on a twinned record; on masked or redacted data they cannot.

Identification accuracy

Identification accuracy is how reliably Nakato finds the sensitive details in a document — direct identifiers such as names, and the indirectly identifying information that pattern matching misses — across 50+ entity types. Higher detection means less leakage. Anything below the 95% confidence gate is held back rather than passed through, so the layer fails closed.

Reversal accuracy

Reversal accuracy is whether every substitution is put back exactly when the answer returns, so the answer you read is about the real customer. Every substitution is also written to the audit trail, so any answer can be traced back through the swaps that produced it.

Leak rate under active extraction attack

The leak rate is the share of cases in which an active extraction attack recovers an original value. A lower rate is better. Under the same attack, GLiNER, an open-source entity recogniser, has a leak rate of 7.8% and Presidio, the open-source PII toolkit, 22%; Nakato’s is on the home page.

Reasoning beyond summarisation

Redaction takes away the values a large language model (LLM) reasons with; Nakato swaps them for semantic twins it can still reason with. On DROP, a benchmark of numerical reasoning over a passage, Nakato holds 99.3% while redaction collapses accuracy to 25% of baseline. On HotpotQA, which asks a model to join facts across documents, Nakato holds 96% and redaction falls to 31.7%. Summarisation, measured above, is one task; these two test reasoning.

Two reasoning benchmarks: Nakato against redaction.
Benchmark Nakato Redaction
DROP Numerical reasoning Holds 99.3% Collapses accuracy to 25% of baseline
HotpotQA Multi-hop reasoning Holds 96% Falls to 31.7%

Reviewing the method

Every figure here is one we will walk you through: the cases, the scoring and the comparators. Write to hello@nakato.ai.