AI / GenAI·6 min·27 August 2026·How I use AI — Part 4 of 4

Blind tasting against the machine: how to test it fairly

A model does not taste. I wrote that in part 1 and it still holds. But underneath it sits a better question: how much of tasting actually lives in language?

If I write down precisely what I observe, and a model gets to the right grape and region from those words alone, then part of my craft is transferable as text. If it doesn’t, I know where the line runs. Both outcomes are worth having, which is exactly why you design the test honestly before you run it.

The job

I want to know what disappears between glass and word. Not whether AI can become a sommelier, which is a boring question with a predictable answer. Rather: which part of a tasting note carries the information, and which part is atmosphere.

For that, the model gets exactly one thing. My words. No photo, no label, no hint about what might have been poured.

The protocol

Six wines, poured by someone else. Bottles out of sight, glasses numbered, order unknown to me.

For each wine I write what I observe in a fixed order: colour, nose, palate, finish. No conclusion, no guess, no region. That is the hardest rule of the whole test, because after two seconds your head wants to attach a name and then goes hunting for evidence to support it.

Then I make my own guess out loud, and it gets written down before the model sees anything. Otherwise you cannot tell afterwards who influenced whom.

Only then does the note go to the model, in an empty conversation, one wine per conversation:

These are my tasting notes on a wine poured blind.
Name the grape, the origin and a vintage range.
Give a confidence from 1 to 5 for each.
Say which detail in my note drove your judgement most.
If the note contains too little to go on, say so instead of guessing.

That last line is not optional. Without it a model always guesses, and then you measure nothing: a lucky guess counts the same as knowledge.

Taken apart

Three choices make this design portable to any field.

Separate observation from conclusion. The note says what you saw, smelled and tasted. Not what you think it is. Skip this and you are testing whether the model can guess your conclusion, which is an easier question that teaches you nothing.

Make the model point at its reasoning. By asking which detail drove the judgement, you see what it leans on. If it turns out to hinge on a single word in five of six wines, you know plenty about the depth of the answer.

Measure confidence separately from correctness. A wrong answer at confidence 2 is a very different thing from a wrong answer at confidence 5. The first is a model that knows its limits; the second is the problem from part 2.

The same three work on an X-ray, a fault report or a legal case. Describe the observation, ask for a judgement with a confidence, and ask what the judgement rests on.

What I expect to happen

This is my prediction, written down before running the test, so I cannot quietly adjust it afterwards.

I think the model will score well on grape for strongly marked varieties, because the standard vocabulary is tightly bound to them. I think it will score badly on vintage, because my notes rarely carry enough signal for that. And I expect that on origin it will lean on one striking term, volcanic for instance, and ignore the rest of the note.

If that holds, it is a dull but usable result. If it doesn’t, that is the article.

Ethics note

There is a difference between learning and judging, and here it matters.

As a learning aid this is lovely. You get language handed to you for what you just observed, and that is exactly the threshold where most people give up on wine. As a judge it is unfit, because it judges a text and not a wine. Confusing the two hands a model authority it cannot carry.

And there is a side to this that is not technical. Judging wine is a trade people earn a living from. When language-model scores start appearing in web shops next to those of real tasters, the question is not whether that can be done, but whether anyone will still explain the difference.

Take it

You do not need six bottles. Three is enough, plus a housemate who pours.

Describe what you observe in four steps: colour, nose, palate, finish.
Write no conclusion.
Then make your own guess and write it down.
Give your note to the model with the prompt above.
Compare three things: who got closer, what the model leaned on,
and how honest it was about its own confidence.

I am running this with six wines shortly and will publish the result, including the times I was wrong myself. That part matters, otherwise it is not a test but a demonstration.

The tasting order I use here is written out at VinoVonk.