Essay · 30. August 2026 · Adrian Verdan
Five AI Reviewers Preferred My Prototype. I Rejected It.

Dieser Bericht erscheint zuerst hier, im englischen Original. Eine Medium-Fassung gibt es noch nicht.
I gave one five-second prototype to five separate AI sessions. Every session preferred it to my old version. Four also found a reason to stop it.
On 28 August 2026, I was making a 55-second video about Kleingewerbe gründen in Deutschland, a downloadable guide for German-speaking beginners starting a small business. The video was meant to show a buyer one manual first step, the practical files inside the guide and where to find the product. It was still a draft; no publishing outlet had been approved, and no release had been authorized.
Drafted with AI assistance; however, every number, failure, and opinion in here is mine
This essay is about the acceptance rule for one scene: an internal decision to use it or reject it. Five AI runs compared two versions of the scene; I checked their objections against the pictures and the guide. None of this measures what real buyers think.
I was testing one five-second animation inside it, not the guide and not the finished video. The animation had a modest job. A person chooses the kind of business they are starting, marks that line on a paper checklist, then opens the same folder's next section: STARTKOSTEN—startup costs.
I wrote a program to draw and animate it. Pages and worksheets from the digital guide appeared on screen as paper in a desk folder; each revision changed how the program drew and moved them.
The old version held the real checklist on screen while a pointer hovered, a checkbox filled and a label changed from OFFEN—OPEN—to ERLEDIGT—DONE. It never connected that decision to the next area of the guide.
My new test version replaced that manual act with a machine-like sequence. A curved arm crossed a shallow desk and pushed cards marked OFFENE AUFGABE—OPEN TASK—toward an archive. Their movement cleared the route to the checklist. It looked purposeful, physical and far more finished than the plain checkbox scene it replaced.
Five fresh AI review sessions chose it. Four of the five also found a flaw serious enough to block it.
I kept both answers.
What I Asked Five AI Sessions to Do
“Reviewers” is shorthand. Each run was one fresh AI session with the same material and no earlier verdict. They were not five people, five independent expert opinions or an audience sample. My notes did not preserve which company's AI model I used, so another person cannot reproduce that part of the test exactly. I should have logged it.
Each run received the same two five-second videos with no sound, labelled A and B, plus six still images from each in time order. I alternated which letter hid the new version. The review instructions required a choice; “neither” was unavailable.
That forced choice answered one narrow question: which of the two supplied versions each run preferred. It could not tell me whether either scene was good enough for a viewer. The old checkbox scored as clear and nearly inert. A landslide against it was useful direction, not a quality certificate.
Each report then scored five things from zero to ten, with half-points allowed: change—how much did the picture change? contact—did an animated object visibly touch what it moved? tension—did the action create resistance or anticipation? originality—was the idea more than a generic checkbox animation? clarity—was the action easy to understand? A separate sentence had to retell the visible chain of cause and effect.
Before the clips were shown, I had written the stop-or-go rule. At least four runs had to prefer the new scene. At least four had to retell its action correctly. Every one of the 25 scores had to reach eight. A Minor was a smaller flaw that did not block the scene. A Major was my label for a flaw serious enough to keep it out of the full video, and the rule allowed none.
Eight was not a scientific boundary. It was a demanding production threshold I chose before seeing the results, meant to stop borderline work rather than estimate audience opinion.
Those thresholds described these five runs only, not 80 percent of any population. The still images could reveal a frame whose visible action contradicted the guide; they could not prove that the animation worked in motion.
The vote was unanimous. The pass was not.
The Winner Failed the Written Rule
Here is the complete first-round record. Each score line follows the order just defined: change, contact, tension, originality and clarity.
Run 1 9 / 7 / 7.5 / 8.5 / 8 Major
Run 2 9 / 8 / 8 / 8 / 7 Major
Run 3 8.5 / 7.5 / 6.5 / 8 / 7 Major
Run 4 9 / 8 / 8 / 8 / 8 Minor
Run 5 8 / 6 / 7 / 7 / 8 Major
Nine of the 25 scores fell below eight. Four reports contained a Major. Only run 4 had all five scores at eight or higher and no Major. All five preferred the prototype and all five retold its broad cause-and-effect sequence, so those two clauses passed. The score floor and Major rule did not.
The reports did not all fail for the same reason. Some saw ambiguous contact. In several frames, pale copies trailed behind the cards like a long-exposure photograph, making solid paper appear to dissolve. The final checklist also seemed to float apart from the desk that was supposed to produce it.
The sequence was visible without a chart. At 0.0 seconds, the cards blocked a locked path. At 2.0 seconds, they overlapped in pale copies while moving toward the archive. By 4.9 seconds, the path was free and a separate checklist had appeared.
The most important objection appeared in one report, not four. Translated from German, it read: “The cards are still labelled OPEN TASK, but enter TASK ARCHIVE and unlock the path; without visible processing, this can read as suppressing unresolved questions rather than clarifying them.”
The report labelled the objection a Major. The frame supplied the check. Its heading said “BLOCKED BY OPEN QUESTIONS.” One visible card asked, “What comes first?” The stack entered an archive anyway, and the path changed from locked to free. The guide required a deliberate decision before the next section opened. The pixels and the intended workflow—not the report's confidence—were why I accepted the objection as a reason to stop.
The scores had already stopped the version, and four reports also broke the no-Major rule. The archive finding showed that one problem lived in the idea, not just its rendering; the written plan for version two removed it alongside the others.
My production note was blunt: “FAIL — METHOD PIVOT REQUIRED.”
Better Than What?
Across the five runs, the new scene averaged 8.7, 7.3, 7.4, 7.9 and 7.6 in the five dimensions, in that order. The old one averaged 2.4, 2.8, 3.1, 3.1 and 9.1. The new scene therefore won on change, contact, tension and originality; the old version won on clarity. Four of the winner's own five averages still sat below eight. The rule was stricter: every individual score, not each average, had to reach eight.

Zum Vergrößern antippen.
A reasonable reader can still object that the comparison paired motion with near stillness and forbade “neither.” The recorded result establishes only that all five sessions preferred the new version to this baseline. The separate acceptance rule asked whether the preferred direction belonged in the film. It did not.
The landslide was real. So was the hill I had chosen.
What the Winner Was Actually Doing
The curved arm approached the cards, but its touch was ambiguous. The cards dissolved toward an archive while still labelled OPEN TASK. Their disappearance released a path. A detached checklist then appeared as the reward.
I had made open questions disappear into an archive and called the result progress.
That reversed the guide's intended workflow. The person was supposed to inspect a real checklist, make a deliberate decision and open the STARTKOSTEN area from the same paper folder. My clip showed a machine removing the inconvenient questions until the path looked clear.
Extra shadow would not repair that sentence. A faster cut would only misstate it with greater confidence.
One AI run could have been enough—if it happened to be the one that raised the archive problem. Four did not raise that problem as a Major. Several fresh passes gave me more chances to surface a missed reading, not consensus. Because I failed to record the model, I cannot know whether those passes shared the same blind spot. Next time I would log the model and use AI systems from different companies, without pretending that made the reports independent. I would reserve that effort for a scene important enough to stop a release.
One AI report did not receive a veto. Evidence did. The reports assigned the labels; I checked each objection against the images and the intended action before acting on it. My notes did not count the suggestions I discarded, so I cannot say how many were wrong. I can make one narrower claim: an AI run surfaced a contradiction I had missed, and the evidence let me confirm it.
Four strong dimensions cannot cancel a false story about cause and effect. “This depicts the wrong operation” deserves a veto.
I Let the Same Idea Win Again
I revised the machine-like desk once more.
Five new runs compared that revision with the old checkbox. In the same order—change, contact, tension, originality and clarity—their score lines were:
Run 1 9 / 9 / 8 / 8 / 9 Pass
Run 2 8 / 8 / 7 / 8 / 8 Score below eight
Run 3 7 / 8 / 5 / 8 / 7 Major
Run 4 9 / 9 / 8 / 9 / 8 Pass
Run 5 8 / 8 / 6 / 8 / 8 Score below eight
Again, all five preferred the desk and could follow the broad action. Only runs 1 and 4 had every score at eight or higher and no Major, so the combined round failed. Every contact score now reached eight, but three tension scores did not. Run 3 scored tension at 5 and filed a Major for a painfully concrete reason: the mechanism's longer arm lay across the checklist sentence the scene was meant to prove.

Zum Vergrößern antippen.
A separate rule, recorded in advance, set a limit of two versions of the same visual idea. My notes treated the first and second desk attempts as those two. Both had failed. The next attempt could use no arm, no archive and no machine. I had to change the concept, object and setting.
Polish had run out of authority.
The Pencil Was a Different Answer
The replacement was almost embarrassingly literal.
The camera looked straight down at a paper-covered desk. The checklist remained in place. A pencil touched its checkbox and stayed there while the mark formed. It lifted, leaving a visible gap, then touched a fold in the same folder. The STARTKOSTEN area opened from that fold.
The new story was simple enough to point at: a person makes a decision; the same document reveals the next piece of work.

Zum Vergrößern antippen.
A third set of five fresh runs compared the paper scene with the old checkbox. All five preferred it and retold the action correctly. Using the same five-score order, this is the full record:
Run 1 9 / 9 / 8 / 9 / 9 Pass
Run 2 9 / 9 / 8 / 8 / 9 Pass
Run 3 9 / 8 / 8 / 9 / 8 Pass
Run 4 9 / 9 / 8 / 9 / 9 Pass
Run 5 9 / 9 / 8 / 9 / 9 Pass
All 25 scores reached eight, and the five reports contained no Major. Every run put tension exactly at the floor of eight, so I treated that as a warning, not a safety margin. The instructions also showed the reviewers that eight was the minimum, so the visible threshold—not just the scene—may have influenced those identical tension scores. Next time, I would collect the scores before showing the threshold. The narrow still-image test passed.
That result carries an obvious bias: a row of still frames may favor the idea that reads most literally when frozen. The paper scene may have benefited. No beginner had evaluated it, and the test said nothing about whether viewers would understand or buy the guide. Its benefit was internal: the sampled pictures no longer contained the contradiction that had stopped the desk scene.
Later, the 55-second file passed automated video and audio checks. My notes still contained no verified start-to-finish viewing with sound, and no external release had been approved. The still-image test had earned the next review, nothing more.
The AI sessions gave me possible readings. Their preference showed which direction looked more alive. The written veto stopped that preference from becoming permission.
I still like the desk scene. That is exactly why I needed the rule.
If you want to separate the AI that builds work from the AI that reviews it, Orchestrating AI Agents contains the role separation and review sequence I use.