How It Works
Your data stays home. Only its twin travels.
Organisations have had to choose between using and exposing their data to AI, or cleansing it and making it unusable. Nakato offers a new option. It sits between a company and any AI model, detecting sensitive data and converting it into fictional twins that preserve context. Only the twin leaves your environment, your data remaining in your domain.
The problem
The models are outside your company, but the data can't travel.
The Large Language Models (LLMs) you want to use run in data centres and are reached over the internet through a third party’s service. A model small enough to run on your own hardware is possible, but has much diminished capabilities.
So a company that wants a model to work with its customer data meets a wall. If that data is sensitive, containing Personally Identifiable Information (PII), letting a third party see it would breach regulation and the responsibilities the data is held under. That applies to almost every business that keeps personal data of some kind.
A model needs context to reason, but with sensitive data, the context is what is sensitive.
The detail that identifies a customer is the same detail that a model reasons with.
The workarounds
Why existing methods keep the privacy but lose the answer
The common solutions today simply remove sensitive details. Redaction blocks out each bit of detected PII. Masking might insert a label for the type of detail, so every name becomes the same word for a name. Tokenisation goes a step further: each value gets its own placeholder, so the first customer is Person 1 wherever they appear. All three protect the data, but don't leave the model anything to reason with.
Take a vulnerability flagging exercise. Joan Mercer, aged 81, living in Whitby, was widowed in March and is calling about a blocked transfer of £4,200 to a builder. Redacted, every identifying detail is blotted out. Masked, the model reads labels where Joan's name, age and blocked transfer amount were. Tokenised, it knows there is one customer with one age, one town and one life event, but not what any of them are: not that she is 81, nor that being widowed in March marks her as vulnerable. In each case the model cannot give a tailored answer, because the context has been removed.
Two things are lost
- Information. Replace a name with a label and you lose any distinction of this case, or that she is a woman, which you may have needed.
- The way back. Redaction and masking cannot be undone, so across a large dataset the results are hard to differentiate and harder to build on. Tokenisation can be reversed, but only after the model has answered without context.
Three steps
Reversible Semantic Pseudonymisation
Each word in the name is one step.
Semantic
Nakato finds sensitive details, direct or indirect, and works out what each one means in context. It reads the task too: its requirements, what details need to be preserved, and what isn't needed.
Pseudonymisation
Each identifying detail is replaced with a twin of equivalent context: fictional, but behaving like the real thing. The twin is what the model sees.
Reversible
Every substitution is recorded locally, in an auditable log. When the answer comes back the original values are swapped back on-premises, and the user sees a complete answer.
One record, three steps
01Semantic
Call from Joan Mercer, 81, of Whitby, widowed in March, about a blocked transfer of £4,200.
In your environmentFound
02Pseudonymisation
Call from Edith Hargreaves, 83, of Filey, widowed in April, about a blocked transfer of £4,050.
What the model seesTwinned
03Reversible
The model’s summary and its flag that she may need extra care, mapped back onto Joan Mercer before anyone reads it.
Back in your environmentRestored
Joan Mercer's record at each step. The model only ever reads the middle one; the answer is mapped back inside your environment.
The twin
Keeping the features the task needs
Any value can be described by a list of its features. A name like Adam is a forename, a man’s name, four letters long, beginning with A. Keep going and the list does not end: Hebrew in origin, biblical, how popular it was the year a customer was born.
A good twin keeps the features the task depends on and changes the rest. Counting the men on a payroll needs a man’s name; calculating vulnerability requires an age. Working out which features matter, for each task, is the semantic part of our process.
This is also the hard part. The features a task might need are open-ended, and they have to be found on your own infrastructure: the obvious tool to locate this, an LLM, is the one place the data cannot go.
Record · Customer call
| Field | Originalstays with you | Twinwhat the model sees | Preserved |
|---|---|---|---|
| Customer | Joan Mercer | Edith Hargreaves | a woman's name |
| Age | 81 | 83 | age band, over 80 |
| Town | Whitby | Filey | small coastal town, same region |
| Life event | widowed in March | widowed in April | recently bereaved |
| Transfer amount | £4,200 | £4,050 | same bracket |
Deployment
It runs inside your domain
Nakato is deployed in your VPC or on your own premises, as a layer between your data and the AI tools you use. Only the twin crosses to the model; the raw data, the mapping of twins and the audit trail stay with you.
- It works with any model you already use.
- Any detail the detector is not certain about is held back, never passed through.
- Every substitution and its reversal is logged, for supervisory review.
Where each thing lives
Stays in your environment
The record, as written
Joan Mercer
The map from each twin to its original
Edith Hargreaves → Joan Mercer
The audit trail of every swap and its reversal
Your boundary
Crosses to the model
The twin, in place of the record
Edith Hargreaves
The task you asked it to do
What stays in your environment, and the only things that cross to the model.
The evidence
Two measures: utility and privacy
Utility is how close the model’s answer, after twinning and reversal, comes to the answer it would have given on the raw data. With the context preserved, it can reason closer to how it would on the original.
Privacy is how reliably sensitive details are found, and how little can be re-identified, even under active attack.
Both are measured, not asserted. See the benchmarks.
The same line, twice: equivalent in every way that matters, different in the one that counts.