Questions

Last updated

What teams ask before they put sensitive data to work with an AI model, answered plainly. The answers on the law link to the ruling and the guidance they rest on, and the answers on accuracy to the benchmarks. The terms are defined in the glossary.

What is Nakato?

Nakato lets regulated businesses put their real data to work with AI. A private layer between sensitive data and the model, it swaps every sensitive detail for a semantic twin at the moment of each query — fictional, but faithful in everything the task depends on — and puts the real details back in the answer. Your real data never leaves your environment.

What is a semantic twin?

A semantic twin is a value that stands in for a real one: fictional, but statistically faithful — same shape, same context, different person. Nakato makes one for each sensitive value at the moment of each query, keeping what the task depends on, such as an age band or a balance to within 2%, and changing who it is about. The model only ever sees the twin.

How does Nakato work?

Nakato works in three steps, which give its method its name: Reversible Semantic Pseudonymisation (RSP). Semantic: it reads the whole document for meaning and finds every sensitive detail, direct or indirect. Pseudonymisation: it swaps each one for a semantic twin, and the model only ever sees the twin. Reversible: when the answer comes back, every swap is exactly reversed and recorded. It happens live, at the moment of each query — not as an up-front dataset transformation. The landing page walks one record through all three steps.

How is a semantic twin different from redaction, masking or tokenisation?

Redaction, masking and tokenisation take the meaning out with the identity; a semantic twin changes the identity and keeps the meaning. Redaction leaves a gap, masking a label and tokenisation a code, and a model cannot reason with any of them. Nakato hands the model a complete, plausible value instead, so it works on a whole record, and puts the real values back in its answer. The benchmarks measure what the difference does to accuracy.

What each approach hands the model, and what it costs the answer.
Approach What the model receives What it costs the answer
Redaction or masking A gap, or a label The context it reasons with
Tokenisation A code with no meaning to read The ages, scales and relationships in the data
A random fake value The right type, picked at random The details the task depended on
A synthetic dataset A fictional dataset, made up front Fidelity to today’s data
On-premise hosting The real data, on a model you run The frontier models
Reactive guardrails Whatever was not blocked Context, without warning
A semantic twin A faithful fiction, made for the query Nothing the task depends on

How is a semantic twin different from a synthetic dataset?

A synthetic dataset is made once, up front, and stays as it was made, so it drifts from the data it imitates and tests run on it do not predict production. Nakato makes a semantic twin live, at the moment of each query, from the real record in front of it. The model reasons on today’s data, faithful in shape and context, and the answer comes back with the real values restored.

Does our real data leave our environment?

No. Your real data never leaves your environment. Nakato runs inside your VPC or on-premise, on a self-hosted small language model, with no dependency on a big API provider. What goes to the model is the twin — a substituted version of each prompt — and the mapping that reverses the swap stays with you. Nothing about your existing data classification has to change. The assurance page has the particulars.

Which AI models does Nakato work with?

Nakato works with any large language model (LLM), because it sits in front of the model, not inside it: Claude, GPT, Gemini, Mistral and fine-tuned models alike. You can switch providers without rewriting your governance. Whichever model you use only ever receives the twin, and the real values come back in its answer, inside your environment.

What happens when Nakato’s detection is unsure?

Nakato fails closed. An entity the detector is not certain about — anything below the 95% confidence gate — is held back rather than passed through to the model. Nakato is designed so the model has no reasonable means of identifying your customers, and a doubtful detection is never a reason to let a real value through.

Does Nakato find indirect identifiers?

Yes. Nakato reads for meaning rather than matching patterns, so it finds indirectly identifying information — the only part-time analyst on a small team, a date that narrows a group to one person — across 50+ entity types. For documents, Holistic Contextual PII Detection (HCPD) goes further: it twins the contextual fingerprints that identify an author without a single name, such as a jurisdiction, a team size, a pilot quarter or a house term.

Is data twinned by Nakato still personal data?

For you, yes: you hold the originals and the mapping that reverses the swap, so it stays personal data in your hands under the GDPR. What can change is the model provider’s position. In EDPS v SRB (C-413/23 P, 4 September 2025), the Court of Justice of the EU held that pseudonymised data is not personal data for every recipient: it depends on whether the recipient has means reasonably likely to be used to identify anyone. Nakato is designed so the model has no reasonable means of identifying your customers.

What is a pseudonymisation domain, and where does the model sit?

A pseudonymisation domain is the environment in which pseudonymised data must not be attributed to a specific person — the European Data Protection Board’s term, from its Guidelines 01/2025 on Pseudonymisation, adopted on 16 January 2025. With Nakato, the model works inside it and only ever sees the twin. The mapping that reverses the swap stays outside it, with you: by the EDPB’s definition, whoever holds that information is outside the domain.

Can we evidence what Nakato substituted?

Yes. Nakato writes every substitution to an audit trail — what was replaced, with what, and when — and records its reversal. A reviewer can retrace any answer through the swaps that produced it without ever seeing the underlying data, and the trail is ready for supervisory review.

Does our data classification have to change?

No. Nothing about your existing data classification has to change. Restricted data stays restricted and stays where it is: Nakato runs inside your environment, and the model only ever receives the twin. The rules you already apply to your data keep applying, and every substitution is on the record for the people who sign them off.

How do you know the twin keeps answers accurate?

Nakato measures the answer, not only the detection: how correct the model’s answer is after pseudonymisation and reversal, scored against its answer on the raw data, across 500,000 financial-services test cases. On two established reasoning benchmarks, tested separately, Nakato holds 99.3% on DROP (numerical reasoning), where redaction collapses accuracy to 25% of baseline, and 96% on HotpotQA (multi-hop reasoning), where redaction falls to 31.7%. The benchmarks page explains each measure, and the four headline figures are on the home page.

Can an AI vendor build Nakato into its product?

Yes — that is what the Reversible Semantic Pseudonymisation (RSP) SDK is for. An AI company whose product falls short on the cleansed data regulated buyers allow can put the layer inside its product: the model only ever meets the twin, its users get the real answer back, and every substitution is on the record. The RSP SDK is in development, and early access is by registration on the developers page.

Who is Nakato for?

Nakato is for regulated businesses that hold data they cannot yet use with AI — teams whose pilots stall on data rules rather than on technology. It sells into financial services first, and it is built for any records that must stay private: insurance, healthcare, legal and HR work the same way, because Nakato cares about data, not sectors. AI companies selling into those businesses can build it into their own products.

Is Nakato hiring?

Yes. Nakato is hiring engineers for a problem most companies cannot even attempt. Generating a twin that keeps its meaning is not a solved problem: it takes an unusual mix of semantic ontology and computer science, and no two domains ask the same thing of it. Write to hello@nakato.ai with what you would bring, and read what the work involves.