Essays / Der Richter schrieb „disconnected“

Essay · 23. September 2026 · Adrian Verdan

The Judge Wrote "Disconnected" and Let It Through

Links: ein Render gestapelter betongrauer Blöcke mit einer limettengrünen Fläche auf schwarzem Raster, ein weißer Kreis um ein kleines graues Teil neben der grünen Fläche und eine runde Lupe, die zeigt, dass es absteht und mit nichts verbunden ist. Rechts: drei Balken, beschriftet matches the order 0.82, follows the style 0.85, free of defects 0.70, jeweils über eine blaue Schwellenmarke hinaus. Der Render links ist eine KI-generierte Illustration.
Das Bild, das mein Bot am 26. Juli gepostet hat, eine Lupe auf dem Teil, das sein neuer Richter später „disconnected“ nannte, und die drei Noten dieses Richters gegen meine Schwellen. Der Render ist eine KI-generierte Illustration meines Bots (FLUX.2-dev), unbearbeitet als Beleg gezeigt; die Komposition darum ist nicht KI-generiert.

Dieser Bericht erschien zuerst hier, im englischen Original. Er steht auch bei Medium, in der Publikation AI Advances Briefs. Dieselbe Fassung bei Medium →

Three AI models judged the images my social-media bot wanted to post. One praised nearly all of them, one gave nearly everything the same grade, and the one I kept wrote the defect into its notes and scored the image above my line anyway.

Drafted with AI assistance; however, every number, failure, and opinion in here is mine

My social-media bot has an image judge: an AI model that looks at every picture before the bot may post it. About the picture at the top of this essay, a stack of grey blocks with one lime-green face, it wrote: "small floating gray triangular facet near top-right edge of accent wedge, disconnected/ambiguous from main mass." In plain words: a small grey piece sticks out beside the lime face and belongs to none of the blocks. In the same answer it scored the picture 0.70 out of 1 for being free of visible defects. On that score my bot now accepts anything from 0.50 up.

Some orientation. I run a one-person company, and one of its AI agents is that bot. It posts to Bluesky and Mastodon several times a day without me. Many of its own posts carry an illustration in one house style: a single abstract object in concrete grey with one lime-green face, on a black field. An image model called FLUX draws it. A second model, the judge, reads the drawing and hands back three scores between 0 and 1, plus a few lines of notes. If all three scores clear my thresholds, the image goes out under my brand, and no person looks at it first. That is the design. This essay is about the judge it rests on.

What a Judge Hands Back

Here is that verdict: all three scores, and three of the four defect notes. The fourth said the lime piece "reads ambiguous rather than a clean solid block".

motif_fidelity   0.82
  (object as ordered?)
brand_fit        0.85
  (follows the style?)
artifact_free    0.70
  (free of defects?)
observed_defects:
- small floating gray
  triangular facet near
  top-right edge of accent
  wedge, disconnected/
  ambiguous from main mass
- heavy uneven grain/
  speckle texture across
  block faces
- shadow under small cube
  blends muddily into
  block behind it

The order was "a stack of solid geometric blocks". Once you have seen the loose grey piece, the whole stack reads as unfinished. The judge found the right words for it, then wrote a number that cleared my line.

This judge is the new one, and it scored this image on 23 September. The image itself is older. My bot posted it on Mastodon on 26 July, under a different judge.

Seven Weeks of "Perfectly"

A bet before the numbers. The old judge, the one my bot used from late July until it died in mid-September, had graded 61 of the images in this essay. How many do you think got the identical grade: 1.0 for matching the order, 0.9 for style, 1.0 for being free of defects?

Twenty-three. More than a third. Fifty-seven of the 61 got 0.9 or better for being free of defects. For 50 of them its defect list was empty, and for 34 it raised no complaint of any kind. The word "perfectly" appears in 20 of its summaries. On the image at the top it wrote: "The image perfectly captures the requested motif of a stack of solid geometric blocks."

The old judge was gemma, an open model from Google that a provider ran for me. To say whether it was right I need a truth, and here is the weakest joint of this essay: my truth is also an AI. It is another of my agents, Claude Opus with my written style rules, which I use as a second pair of eyes on much of what goes out under my name, though never as a gate on the bot's images: it labels test sets and looks at what went out afterwards. In July it labeled 26 images blind, without seeing any judge's verdict. In August it labeled 12 more, and that time it knew which ones had already been published. Together that is 19 images it would publish, 17 rejects and two it called borderline; with 24 more images it never saw, my test set has 62. I did not label them myself, and I do not let the reviewer judge live images: the first would put me back in the loop the bot exists to remove, and a gate that needs me is a gate I skip on a busy day; the second would leave nothing independent to measure the judge against. Even so, the reviewer and the judge I ended up with are cousins, both Claude; if they share a blind spot, my numbers count it as agreement, and nothing in this essay can rule that out.

On the 17 rejects, the old judge's thresholds let 9 through. For 15 of the 17 its defect list was empty. Five of the 17 had in fact been posted, the image at the top among them. Another was a double gate in a concrete frame whose left wing dissolves into loose slats and black air, which went out on Bluesky on 23 August.

Then, on 16 September, the old judge stopped answering. Its provider refused every request: the model was no longer available. Between 16 and 22 September my bot paid for 41 images and got 0 verdicts. The money is the smallest part, 63 cents. The technical log said it plainly, once per image: "HTTP Error 400", a refused request. But the line that travelled upward, into the audit trail and the daily report the bot writes for me, read "all rendered images failed the quality gate." Up there, a dead judge and a very strict judge write the same line. Since 23 September a missing verdict has its own line.

The daily report noticed the missing images from 16 September and on its first two days even gave a reason, the wrong one. In seven days of reports the judge's name does not appear once. The last image my bot published went out on 14 September.

A Judge With One Grade

My first replacement candidate was Gemini Flash, which I tried on a handful of recent images, none of them labeled by my reviewer. The nine it answered had grades from the old judge between 0.40 and 1.0 for matching the order. Gemini's scores ran from 0.52 to 0.84, and on being free of defects every answer sat between 0.72 and 0.84. An image the old judge had scored 0.40 for matching the order got 0.63 and 0.68 in two tries; one it had scored 1.0 got 0.76. Under the thresholds of the time, which asked for 0.90 on defects, Gemini approved nothing. That looks strict. It is the opposite of judging: nearly the same number for everything.

The obvious fix is to explain the scale, so I did, on six images with a slightly older Flash version: 1.0 means no visible defect, and a clean render must score 0.95 to 1.0. The average score moved by less than two hundredths, and all 36 numbers, with and without the explanation, landed between 0.68 and 0.78.

That was the most useful morning of the week, because it changed the question I ask first. Before "is the judge right?" comes "does it spread at all?" A judge that gives one grade to images that clearly differ is a constant with a delay. You can recalibrate a judge whose scale is shifted, but a constant can only be replaced. The test costs next to nothing, and it needs no truth yet, only images that differ. The authors of G-Eval (Liu and others, 2023) saw a milder form of this when models grade text: one score swallows most of the scale.

The second candidate was Claude Sonnet, the same model family as my reviewer, called through the same command-line tool I use for everything else. Same instruction as the old judge, same 62 images. It spread: from 0.05 to 0.97 for matching the order, from 0.30 to 0.95 for being free of defects.

Dot chart in three rows on a scale from 0 to 1 for the score free of visible defects. Old judge, 61 images: almost all dots stacked at 0.9 and 1.0, including orange dots for images my reviewer would not publish. Gemini Flash, 13 answers on 9 unlabelled images, values rounded to 0.05: a narrow cluster between 0.72 and 0.84. Claude Sonnet, 62 images: dots spread from 0.30 to 0.95, orange dots mostly lower, teal dots mostly higher.

Zum Vergrößern antippen.

Each dot is one image or answer, placed by the score a judge gave it for being free of defects. The old judge (top, 61 images) piles up at 0.9 and 1.0, orange dots included. Gemini Flash (middle) saw only nine images, four of them also in the other two rows, none labeled by my reviewer, and sits in a narrow band. Grey means no label or borderline; values are rounded to the nearest 0.05. Claude Sonnet (bottom, 62 images) spreads out, and the images my reviewer would not publish (orange) sit lower on average, though not all of them. Data: my bot's calibration file and trial logs, 22 and 23 September 2026.

The Price of Zero

Here is the plainest test I know for being right. Take one image my reviewer would publish and one it would not, and ask which one the judge scored higher for being free of defects. Over all 323 such pairs, the new judge orders the pair the way my reviewer does 8 times in 10. The old judge got 5 in 10, a coin toss, and for half of its pairs it gave both images the same score; on the 306 pairs both judges saw, the result is the same.

When I opened this task I had written down what I expected: thresholds that let 40 to 60 percent of the 62 images through while accepting none of the 17 rejects. And an abort line: if every set of thresholds that lets at least 30 percent through also lets a reject through, I do not switch.

I tried every combination of the three thresholds from 0.40 to 0.95 in steps of 0.05, 1,728 of them. For each number of rejects I allow, one combination lets the most images through; drawn as a line, that is the price of each reject: each of the first three buys five to ten more images, and after that the line flattens. None met the expectation. The best combination with no reject let 4 of the 62 through, 6.5 percent: one image my reviewer had approved and three it had never seen. The old judge, strangely, did a little better there, 6 of its 61. At 40 percent, the fewest rejects I could get was three. Zero rejects would have starved the bot of images, and the abort line was crossed.

Line chart: horizontal axis rejects that get through, 0 to 11 of 17; vertical axis share of images that pass. A black line for the new judge rises from 6.5 percent at zero rejects to 40 percent at three and about 63 percent at eight. An orange step line for the old judge stays at about 10 percent until seven rejects, then jumps to 52 percent at eight and 62 percent at nine. A shaded box at zero rejects and 40 to 60 percent, labelled what I asked for, is empty. Marked: my thresholds, 25 pass with 3 rejects; thresholds plus reading the notes, 20 pass with 1 reject; old judge as it ran, 38 of 61 with 9 rejects.

Zum Vergrößern antippen.

For each number of my reviewer's 17 rejects allowed through (left to right), the largest share of images any threshold combination lets pass. Black: the new judge, orange: the old judge, 1,728 combinations each. With zero rejects the new judge passes 6.5 percent; below eight rejects the old judge passes at most 6 of 61 images. The blue box, what I had asked for, stays empty. Data: my bot's calibration file, 23 September 2026.

So I broke my own rule. I switched anyway, with thresholds of 0.60 for matching the order, 0.70 for style and 0.50 for defects, which pass 25 of the 62 and let 3 of the 17 rejects through, where the old judge, passing 38 of its 61, let 9 through. The decision record says that I switched against my own abort line, and why. At the old judge's own pass rate the two are close, 8 rejects through for the new one against 9. Below that they part: to let fewer than 8 rejects through, the old judge had to shut out all but 6 of its 61 images, and the new one keeps 40 percent with 3. And no judge meant no images at all, while my own brief for the replacement had asked for checks that are sensible but not too strict, and for enough images to post. The abort line did not stop me, and I do not think it should have. Without one I would have made the same switch and felt careful about it. With one, the switch had to be argued in writing, and the argument had to name its price: three rejects among the 17.

Read the Notes

Then I looked at what the three rejects had in common and found the grey piece again. The new judge had not missed the defects. On two of the three it had described them: "disconnected" on the image at the top, and on another "an extra angled accent-color wedge/panel is fused into the block". It had seen them, written them down and scored past them.

Rejecting every image with any defect note would reject all 62; this judge always finds something. So the new rule reads for particular words. If the judge's own notes contain one of a short list of terms, the image is rejected whatever the scores say. The list: disconnected, detached, unattached, fused and its forms, protrude and its forms, not seated, unseated. I tested it on the 62 stored verdicts before switching it on. Of the 25 images that cleared the thresholds it removed five: two rejects, one image my reviewer would publish ("some pieces appear fused/overlapping rather than simply stacked") and two it had never labeled. With the rule, 20 of 62 pass, 32 percent, and one reject gets through. Thresholds alone, allowing that same single reject, pass at most 14.

Six renders of grey-and-lime block objects in a two-by-three grid, each with my reviewer's label, the new judge's three scores and a quote from its notes. Top row: two rejects whose notes contain disconnected and fused, stopped by the word rule. Middle row: an image my reviewer would publish whose note contains fused, stopped too as the price of the rule; a tilted block, reject, whose notes mention only grain, gets through, posted on 26 August. Bottom row: a clean pendulum image that passes everything; a double gate with a dissolving wing, stopped by the scores, which the old judge had scored 1.0 for defects and posted on 23 August.

Zum Vergrößern antippen.

Six images with the new judge's notes quoted exactly. The top four clear all three thresholds: the word rule stops the first two, which my reviewer rejected, and the third, which my reviewer would have published; the fourth, a tilted block, gets through, and my bot had already posted it on 26 August, under the old judge. The bottom two, for comparison: a clean image that passes everything, and the dissolving gate, which the new judge's scores alone stop. Renders: AI-generated (FLUX.2-dev), unretouched, cropped. Scores and notes: verbatim model output. Labels: my AI brand reviewer.

Two words are missing from the list on purpose. "Floating" and "gap" would also catch defects, but the judge uses them for good images too: an object that hovers in the black void by design, or sits in a shadowed notch. Adding them would have rejected three images my reviewer passed and one it called borderline. A word list is crude. It cannot read "no fused edges", and it will need measuring again whenever the judge's vocabulary drifts. It is also the cheapest check I have: it reads text the judge has already written and costs no extra call.

I also tried the opposite of reading the notes: a paragraph in the instruction that told the judge, check by check, what to punish, nothing floating and nothing poking out among them. On the pair test for defects it fell from 8 in 10 to 6 in 10, and with three rejects allowed it let 20 images through instead of 25. Being told harder made it worse at telling good from bad. Only with zero rejects allowed did it pass more, 9 images instead of 4, still far from what I had asked for.

My reading, and it is a reading rather than a measurement of the model's insides: the score is a summary the model writes after it has looked, and summaries forgive. The notes are closer to what it saw. Chiang and Lee (2023) found a neighbouring effect with text: asking a judge to explain its rating improves its agreement with people. I had been asking my judge for the explanation all along and throwing it away.

What Still Gets Through

The reject the rule misses is a lime-and-grey block tilted in mid-air, whose faces do not close into a solid body. It has no floor and no shadow. My bot posted it on Mastodon on 26 August, under the old judge. The new judge gives it 0.90 for the order, 0.92 for style and 0.85 for defects, and writes about grain, grid spacing and the angle of the tilt. Nothing in its notes for the list to find.

The first two test runs of the new judge in the bot's real pipeline, on 23 September, ran against a copy of my bot's database, so nothing could be posted. Both images had a fault, and the setup stopped neither. Both orders were "a large concrete cube with one corner cut away, a smaller cube seated in the notch". In the first image the small cube pokes out past the edge of the large one. The judge gave it 0.95 for matching the order and wrote about a glow and some grid lines; my reviewer caught the fault afterwards. The second image, tested with the word rule already on, carries a lime triangle on the side of the cube that nobody ordered. This time the judge saw it and wrote: "Unrequested diagonal triangular accent cutout on the right face of the cube, not part of the described single-notch/small-cube motif." Then it scored the image 0.65 for matching the order, five hundredths above my line. None of my words were in that sentence, so the image passed. "Unrequested" is not on my list, and adding it needs the same measuring as the others.

That is where I stand. Measured on the same images I used to choose the rule, and against a reviewer that is itself a model, the new setup lets 1 of the 17 known rejects through where the old one let 9. That is a lab figure, the best case, and the two test runs above are what the real pipeline has shown so far. Of the five rejects my bot had actually posted, the new rule would have stopped four, the image at the top and the dissolving gate among them. The tilted block would still be out there. The rule has an abort line of its own, written before the results: my reviewer looks at the first ten images the bot publishes under it, and if more than two of them show a visible defect of this kind, the rule gets stricter. Once ten have gone out, I will know. What I trust more than the scores now are the judge's sentences. What I do not trust yet is my list for reading them: it has a hole the size of every word I have not thought of, and the last test walked straight through it.

How I Counted

The 62 images are my bot's calibration set: 26 labeled blind by my reviewer in July, 12 labeled in August with knowledge of what had been published, and 24 drawn at random from late August to mid-September, half of them images the old judge had approved and half ones it had rejected; those 24 carry no label from my reviewer. The old judge graded 61 of them. The new judge's scores are one call per image with the instruction the bot uses in production, stored with the notes. The pair test counts every pair of one image my reviewer would publish and one it would not (323 pairs for the new judge, 306 for the old one), with ties counted as half. As rank correlations against my reviewer's 36 publish-or-not labels, where 0 means no relation and 1 the same order, the new judge reaches 0.40 for matching the order, 0.50 for style and 0.53 for defects; for comparison, the VIEScore paper (Ku and others, 2024) reports about 0.4 for GPT-4o against human raters on generated images, where two humans agree at about 0.45, though that is against people and mine is against a model. The Gemini figures come from 22 September: two runs of Gemini 3.8 Flash with 40 requests, of which 13 answered, covering nine images outside my reviewer's labels, and an A/B test of six images on Gemini 3.6 Flash. Images and reviews per day come from my bot's cost ledger. The thresholds were chosen on the same 62 images they are evaluated on, so every pass rate here is optimistic. All counts were taken on 23 September 2026 from the files named above.


The same trap, a green number standing in for a question nobody asked, runs through Nobody Opens a Passing Check, my book of thirty-nine logged days with two AI agents.

Weiterlesen

Alle Berichte im Überblick.

Jeder Bericht geht von einem konkreten Fehlschlag aus, mit Datum, Logauszug und der Änderung, die daraus folgte: zur Essay-Übersicht →