Glossary

Last updated

The words Nakato uses, and the regulator’s words it works with, each defined in a sentence you can take on its own. Nakato’s terms come first, then the regulator’s, then the measures behind the benchmarks.

Semantic twin

A semantic twin is a value that stands in for a real one: fictional, but statistically faithful — same shape, same context, different person. Nakato makes one for each sensitive value at the moment of each query, keeping everything the task depends on and changing who the value is about. The model only ever sees the twin.

Reversible Semantic Pseudonymisation (RSP)

Reversible Semantic Pseudonymisation (RSP) is Nakato’s method for putting sensitive data to work with an AI model: read the record for meaning, replace anything that identifies a person with a semantic twin, and put the real thing back in the answer. It happens live, at the moment of each query — not as an up-front dataset transformation. The landing page walks one record through all three steps.

Reversal

Reversal is the last step of Reversible Semantic Pseudonymisation: when the model’s answer comes back, Nakato swaps every twin for the original value it stood in for, so the answer is about the real customer. Other tools call this re-identification or rehydration. Only the people already authorised to see the original ever see both sides of the swap.

Holistic Contextual PII Detection (HCPD)

Holistic Contextual PII Detection (HCPD) is Nakato’s extension of the twin from names to the contextual fingerprints that identify a document’s author — a jurisdiction, a team size, a pilot quarter, a house term. It is built for legal agreements, board papers and investor materials, which can give their author away without a single name. The HCPD page steps through one such document.

Contextual fingerprint

A contextual fingerprint is a detail in a document that points to who wrote it, or who it is about, without naming anyone — a jurisdiction, an office, the size of a team, a quarter, a house term. Together, a few of them can narrow a reader’s guess to one company. Holistic Contextual PII Detection twins them alongside the names.

Indirect identifier

An indirect identifier, or quasi-identifier, is a detail that names no one on its own but can identify a person in combination with others — a date of birth with a postcode, or the only part-time analyst on a small team. Pattern matching misses it, because there is no pattern to match; Nakato reads for meaning, and twins it too.

Confidence gate

The confidence gate is the certainty, set at 95%, below which Nakato holds an entity back rather than passing it through to the model. It is what makes the layer fail-closed: when detection is unsure, the real value is not sent.

Audit trail

Nakato’s audit trail is the record of every substitution it makes — what was replaced, with what, and when — and of each reversal. A reviewer can retrace any answer through the swaps that produced it without ever seeing the underlying data, and the trail is ready for supervisory review.

Pseudonymisation

Pseudonymisation is the GDPR’s term (Article 4(5)) for processing personal data so that it can no longer be attributed to a specific person without additional information, which is kept separately and protected. Pseudonymised data stays personal data for whoever holds the information that reverses it. Nakato pseudonymises by substitution: each identifying value becomes a semantic twin, and the mapping that reverses the swap stays in your environment.

GDPR, Article 4(5)

“‘pseudonymisation’ means the processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately and is subject to technical and organisational measures to ensure that the personal data are not attributed to an identified or identifiable natural person”

Pseudonymisation domain

A pseudonymisation domain is the environment in which pseudonymised data must not be attributed to a specific person — the European Data Protection Board’s term, from its Guidelines 01/2025 on Pseudonymisation. Whoever holds the information that reverses the pseudonymisation is, by definition, outside it. With Nakato, the model works inside the pseudonymisation domain and only ever sees the twin; the mapping stays outside it, with you.

Means reasonably likely to be used

Means reasonably likely to be used is the GDPR’s test for whether a person is identifiable (Recital 26): every means the controller or anyone else is reasonably likely to use, weighing the cost, the time and the technology available. In EDPS v SRB (C-413/23 P, 4 September 2025), the Court of Justice of the EU applied the same test recipient by recipient: pseudonymised data may not be personal data for a recipient without such means.

GDPR, Recital 26

“To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly.”

Motivated intruder test

The motivated intruder test is how the UK’s Information Commissioner’s Office asks whether someone could be identified from data: could a reasonably competent person with no prior knowledge — using the internet, libraries and public records, but no specialist skills or criminal means — identify them? A document’s contextual fingerprints can fail it on their own, which is what Holistic Contextual PII Detection is designed against.

Downstream utility

Downstream utility is how correct an AI model’s answer is after pseudonymisation and reversal, scored against its answer on the raw data. It measures what a privacy layer costs the work itself, which a detection score alone cannot show. Nakato’s figure is on the home page, and the benchmarks page explains the measure.

Leak rate under active extraction attack

Leak rate under active extraction attack is the share of cases in which a deliberate attempt to recover the original values succeeds. A lower rate is better. Nakato’s figure, beside GLiNER’s and Presidio’s under the same attack, is on the home page; the benchmarks page explains the measure.

Pilot purgatory

Pilot purgatory is where an AI project stalls: it works as a proof of concept but never reaches production. In a regulated business it is often a data problem rather than a technology one: the data the pilot may use has lost the context the model needed. A semantic twin keeps that context, so the work can run on data that behaves like the real thing.