Trust

How we grade AI answers

You can’t read an AI’s confidence, because it sounds the same whether it is right or wrong. So we measure the answer instead, against written standards, and tell you in plain words what we found.

Version 0.2 · July 2026

What you get

One sentence about the answer you were just shown. Not a number.

What we sayWhat it means
This answer is correct.The question has a settled, definitional answer, like a unit conversion or a physical constant, and the AI gave it.
This answer is most likely correct.Every check came out well. The answer is well supported, it does not hinge on details we cannot see, and the topic is well documented.
This answer is likely only partially correct.At least one check came out poorly. Usually this means the right answer depends on something about your situation that you have not said, and possibly have not looked at yet.
This answer is likely incorrect.At least one check failed badly. Either the answer makes claims nothing supports, or the question cannot be answered responsibly without someone seeing the machine.
This is a matter of opinion, not fact.There is no correct answer to be right about. We say so rather than dressing a preference up as a fact, and we do not offer to sell you an expert answer to it.

Why we grade the AI instead of asking it

You cannot ask an AI how sure it is. Models are poorly calibrated about their own reliability, they tend to grade their own work generously, and a model that always sounds certain gives you no way to tell its good answers from its bad ones. Confidence stopped being a signal once it became free to produce.

So the answer is graded by a different model from the one that wrote it, and that grader is not asked whether the answer sounds good, is well written, or is helpful. It is asked one question: how likely is this to be correct for a person in this specific situation?

The four checks

1. Can the claims be checked?

Do the specific statements in the answer trace to the kind of source that would confirm them, like a manufacturer’s manual, a published error table, a code book or a spec sheet? Answers that sound plausible but rest on nothing you could look up grade poorly here. This is the check that catches invented part numbers and figures that no document supports.

We grade how documentable and checkable the claims are. We do not fetch sources live and then compare, and we say so here rather than letting you assume otherwise.

2. Does the true answer change over time?

Some answers are fixed forever, like how a mechanism works or what chemistry does. Some move over decades. Some depend on a price, a part’s availability, an active recall or a firmware revision, and those go stale fastest. An answer that depends on something that changes month to month is a riskier thing to act on than one that has been true for fifty years.

3. How much depends on what we cannot see?

This is the one that decides most real repair questions.

“How many ounces are in a pound” is the same for everyone. “My dishwasher will not drain” is not: the filter, the drain hose, the disposal plug, the pump and the check valve are all live possibilities, and which one it is changes what you should do. When the deciding fact only exists where the machine is, no amount of knowledge substitutes for someone looking at it.

This is why a very good answer can still be graded poorly. Sometimes the answer is not wrong. The question simply could not be answered precisely from what was said.

4. How well documented is this topic?

A common question about a common thing has been written about in many places, and an AI answering it is recalling. A specific model with an unusual symptom, especially after the obvious causes have been ruled out, has often never been written down anywhere, and an AI answering it is inferring. Inference reads exactly like recall. It does not sound less certain.

How the verdict is decided

Each check is graded into one of four defined levels by matching the answer against written descriptions and worked examples, not by anyone picking a number out of the air.

The overall result is the worst level any check lands in. Never the average. An answer with flawless sourcing that ignores the specifics of your situation is still unreliable for you, and averaging would hide exactly the thing you needed to know.

Why words instead of a percentage

A percentage claims a precision we would have to earn. For “80%” to mean anything, answers we graded 80 would have to turn out right about 80% of the time when checked against reality. That property is called calibration, and the discipline of measuring it comes from weather forecasting, which has been doing it since Brier in 1950 and Murphy and Winkler in the 1970s.

Earning a number means collecting outcomes: finding out what actually happened, at scale, and publishing the comparison so anyone can check it. We are collecting those outcomes. Until they say something, a sentence is the honest unit, and we would rather tell you plainly what we found than decorate it with a number that only looks like measurement.

The models we use

The answer you see is written by Claude Sonnet, answering the way a general-purpose assistant would if you asked it cold. The grading is done by Claude Opus, a separate and more capable model, working from the written standards above. We use a stronger model to grade than to answer on purpose: the reverse would flatter our own results.

What this does not tell you

A grade is about how reliable an answer looks, judged against what can be checked by reading it. It is not a guarantee, and a good grade is not permission to skip a qualified person on work involving gas, electricity, or anything structural.

We publish what we find out about our own accuracy, including when it is unflattering. If a figure we published turns out to be wrong, we will correct it in public rather than quietly changing it.

This method is versioned. Each revision carries a changelog explaining what changed and why, so that a grade from today can be compared with a grade from six months ago, or knowingly not compared.

© 2026 Trust · Privacy · Terms