How It Works

Your data stays home. Only its twin travels.

Organisations have had to choose between using and exposing their data to AI, or cleansing it and making it unusable. Nakato offers a new option. It sits between a company and any AI model, detecting sensitive data and converting it into fictional twins that preserve context. Only the twin leaves your environment, your data remaining in your domain.

The problem

The models are outside your company, but the data can't travel.

The Large Language Models (LLMs) you want to use run in data centres and are reached over the internet through a third party’s service. A model small enough to run on your own hardware is possible, but has much diminished capabilities.

So a company that wants a model to work with its customer data meets a wall. If that data is sensitive, containing Personally Identifiable Information (PII), letting a third party see it would breach regulation and the responsibilities the data is held under. That applies to almost every business that keeps personal data of some kind.

A model needs context to reason, but with sensitive data, the context is what is sensitive.

The detail that identifies a customer is the same detail that a model reasons with.

The workarounds

Why existing methods keep the privacy but lose the answer

The common solutions today simply remove sensitive details. Redaction blocks out each bit of detected PII. Masking might insert a label for the type of detail, so every name becomes the same word for a name. Tokenisation goes a step further: each value gets its own placeholder, so the first customer is Person 1 wherever they appear. All three protect the data, but don't leave the model anything to reason with.

Take a vulnerability flagging exercise. Joan Mercer, aged 81, living in Whitby, was widowed in March and is calling about a blocked transfer of £4,200 to a builder. Redacted, every identifying detail is blotted out. Masked, the model reads labels where Joan's name, age and blocked transfer amount were. Tokenised, it knows there is one customer with one age, one town and one life event, but not what any of them are: not that she is 81, nor that being widowed in March marks her as vulnerable. In each case the model cannot give a tailored answer, because the context has been removed.

Two things are lost

  • Information. Replace a name with a label and you lose any distinction of this case, or that she is a woman, which you may have needed.
  • The way back. Redaction and masking cannot be undone, so across a large dataset the results are hard to differentiate and harder to build on. Tokenisation can be reversed, but only after the model has answered without context.

Three steps

Reversible Semantic Pseudonymisation

Each word in the name is one step.

Semantic

Nakato finds sensitive details, direct or indirect, and works out what each one means in context. It reads the task too: its requirements, what details need to be preserved, and what isn't needed.

Pseudonymisation

Each identifying detail is replaced with a twin of equivalent context: fictional, but behaving like the real thing. The twin is what the model sees.

Reversible

Every substitution is recorded locally, in an auditable log. When the answer comes back the original values are swapped back on-premises, and the user sees a complete answer.

One record, three steps

  1. 01Semantic

    Call from Joan Mercer, 81, of Whitby, widowed in March, about a blocked transfer of £4,200.

    In your environmentFound

  2. 02Pseudonymisation

    Call from Edith Hargreaves, 83, of Filey, widowed in April, about a blocked transfer of £4,050.

    What the model seesTwinned

  3. 03Reversible

    The model’s summary and its flag that she may need extra care, mapped back onto Joan Mercer before anyone reads it.

    Back in your environmentRestored

Joan Mercer's record at each step. The model only ever reads the middle one; the answer is mapped back inside your environment.

The twin

Keeping the features the task needs

Any value can be described by a list of its features. A name like Adam is a forename, a man’s name, four letters long, beginning with A. Keep going and the list does not end: Hebrew in origin, biblical, how popular it was the year a customer was born.

A good twin keeps the features the task depends on and changes the rest. Counting the men on a payroll needs a man’s name; calculating vulnerability requires an age. Working out which features matter, for each task, is the semantic part of our process.

This is also the hard part. The features a task might need are open-ended, and they have to be found on your own infrastructure: the obvious tool to locate this, an LLM, is the one place the data cannot go.

Record · Customer call

FieldOriginalstays with youTwinwhat the model seesPreserved
CustomerJoan MercerEdith Hargreavesa woman's name
Age8183age band, over 80
TownWhitbyFileysmall coastal town, same region
Life eventwidowed in Marchwidowed in Aprilrecently bereaved
Transfer amount£4,200£4,050same bracket
Name and town change; the age, life event and transfer amount stay in equivalent brackets but are materially different. The model reasons on a record that behaves the same.

Deployment

It runs inside your domain

Nakato is deployed in your VPC or on your own premises, as a layer between your data and the AI tools you use. Only the twin crosses to the model; the raw data, the mapping of twins and the audit trail stay with you.

  • It works with any model you already use.
  • Any detail the detector is not certain about is held back, never passed through.
  • Every substitution and its reversal is logged, for supervisory review.

Where each thing lives

Stays in your environment

  • The record, as written

    Joan Mercer

  • The map from each twin to its original

    Edith Hargreaves → Joan Mercer

  • The audit trail of every swap and its reversal

Your boundary

Crosses to the model

  • The twin, in place of the record

    Edith Hargreaves

  • The task you asked it to do

What stays in your environment, and the only things that cross to the model.

The evidence

Two measures: utility and privacy

Utility is how close the model’s answer, after twinning and reversal, comes to the answer it would have given on the raw data. With the context preserved, it can reason closer to how it would on the original.

Privacy is how reliably sensitive details are found, and how little can be re-identified, even under active attack.

Both are measured, not asserted. See the benchmarks.

The same line, twice: equivalent in every way that matters, different in the one that counts.