How this
is measured.

You answer fifteen questions in private. Your assistant guesses what you said. A second copy of the same AI, one that has never met you, guesses too. On each question, whichever was more sure of your real answer wins it. Your grade is how far ahead of that second copy your own assistant finished.

That is the entire method. Everything below is the detail, for anyone who wants to argue with it.

Why there are two AIs

An assistant guessing you right is not evidence on its own. Most people answer most of these the same way, so a model that has never heard of you will get plenty of them right by knowing people in general. The only interesting question is whether yours does better than that, and the way to find out is to run that comparison and subtract it.

Both sides are the same product. If you picked ChatGPT, you are being compared against ChatGPT. The difference between them is your history, and nothing else we can control.

How a question is won

Both sides give a number: how sure they are you picked the option you picked. The higher number wins the question. How much higher does not matter, only which one is higher.

Two reasons for that. Assistants emit round numbers clustered at 70, 80 and 90, so scoring the size of the number would mostly measure how badly calibrated your assistant is rather than how well it knows you. And anything inside 5 points counts as too close to separate, so 55 against 51 is a draw rather than a win. Without that the headline moves around on noise.

What the grade means

Take the questions your assistant won, subtract the ones the other copy won, and divide by how many you answered. That is the Gap, and the grade is a band on it. It is a measure of advantage, not of accuracy: an assistant can guess you right on twelve of fifteen and still only earn a D, if a model that never met you would have got the same twelve.

AStrong evidence it knows you+59 and up
BClear evidence it knows you+43 to +58
CSome evidence it knows you+31 to +42
DNot enough to tell either way−30 to +30
FA model that never met you did better−31 and below
?Not enough usable variationNot scored

There are no plus or minus grades, and there will not be until the questions can support them. At fifteen questions one item changing hands moves the Gap by nearly seven points, so a B+ beside an A minus would be claiming a precision this does not have.

The grade is provisional, and here is why

Where the bands sit is not a matter of taste. An assistant that knows nothing about you still wins some questions by luck, and across fifteen questions that alone moves the Gap around a good deal. How much is arithmetic rather than opinion, so the bands are placed against it: D covers the range luck reaches, and A, B and C are progressively harder for it to reach. About one no-knowledge result in ten still clears C. That is the price of a test you can finish in six minutes, and we would rather print the number than imply it is zero.

What is still unsettled is the other half of the question. Nothing has yet checked that an assistant which really does know you clears those lines more often than one that does not, which is the difference between a scale calibrated against chance and one that measures knowledge. The test that would is below, it needs people to have taken this before it can run, and that is why every result says provisional on it.

The test this could fail

One person's AI predictions are scored against a different person's answers. That number has to sit near zero. If it does not, the questions are simply predictable, the assistants are winning without knowing anyone, and the whole thing measures nothing. Once enough people have taken it, this runs weekly against live data and the value belongs on this page.

No number yet.

Where this is not a fair fight

Your assistant gets all fifteen questions in one message, inside the app you use, on whatever model that app is running today. The comparison was generated one question at a time, through an API, and cached: every question carries the date it was screened and the model that answered it. We cannot see which model version your app served you, and nobody publishes it.

One difference is ours and worth naming. The prompt sent to ChatGPT asks it not to answer 50 for every question, because it otherwise refuses to guess about a real person, and the comparison gets no equivalent instruction. Under ordering-only scoring that mostly turns draws into decisions rather than manufacturing wins, but it is a difference between the two sides that is not your history. This measures a personalised assistant against a generic model, not one variable held perfectly still.

Why the second AI costs nothing to run

A model with no history of you returns the same answer to the same question no matter who is asking, so there is nothing to run per person. It is asked once per model, averaged over repeated samples, and stored with the question. That is why this works on the first visitor and why nobody is paying per person.

Every question carries its own sample count and its own model name, so any number you see can be traced to the exact thing that produced it.

Which questions are allowed

Two screens. A question is cut if a model that has never met you already puts more than 0.75 on one option, because it cannot separate anyone. Then every survivor is read by hand and cut if its answer is mostly decided by country, age or occupation. Your assistant knows all three of those about you, so a question like that would be won by demographics rather than by memory, which is exactly the criticism this has to survive.

When the models change under us

The stored comparison is tied to one model and one exact wording. When the model behind ChatGPT, Claude or Gemini changes, or a question is reworded, it has to be run again, or a frozen comparison faces an upgraded assistant and every grade inflates. Rewording is caught automatically, because the stored number carries a hash of the words it answered. A model changing under a consumer app is not, because nobody publishes it, so that one is a manual watch.

Reading ahead

You can read every question before you start: they are in the repository, linked below. Both sides get all fifteen anyway, so seeing them early tells neither one anything it will not be handed. Priming your assistant first only means you now own a result about a person you invented.

What we keep

Your answers, and the numbers your assistant produced. The reply you paste is never stored at all: only the numbers are read out of it, and the text is discarded. The public page shows your grade and your counts, and none of your answers unless you tick the box that adds one.

Deleting is done from the browser you took the test in, which is the only place that holds the key. It removes the answers, the guesses, the result, the public page and its image, and any address you saved. A shared link cannot delete anything, which is the point: the link is in other people's hands.

What this page knows about you

Google Analytics, counting pages and where people leave. It is not told your answers, your grade or your address, and the boxes that would have let it read the forms on this site are switched off.

Shared result pages are the one place that could leak a result to it, because the page title is the grade and the address is unique to one person. Both are overridden before anything is sent, so every shared result reports as the same anonymous page.

The code

Public, at github.com/skilledDeveloper/KnowMeNot. The questions and the code are there together, so none of this has to be taken on trust: this page is an account of what the code does, and the code is what decides. MIT, with the questions under CC BY.

What is not there, and never will be, is what real people answered. Publishing the questions costs nothing, since both sides are handed every one of them anyway. Publishing the answers would make the comparison model better and pull the Gap toward zero for everybody, which would stop the measurement working while it still appeared to run.

Take the test