Knowledge / What the pipeline rejected.

Knowledge · Free · AI-assisted · Status July 2026

What the pipeline rejected.

One run, four agent roles, one context-separated cross-check. Across the eleven checked key claims: no contradiction finding — and still nine interventions: four corrections, two rejected image candidates, one preserved ambiguity, one claim that was stopped and later substantiated, one intervention on the tool itself.

The task

One paper, four agent roles, one human at the end.

On 20 July 2026, an astrophysics paper went through an internal pipeline of four agent roles: research, writing, visualization, and a context-separated cross-check with its own research. The subject was Tzanidakis & Davenport (2026), published in March 2026 as arXiv:2603.10952 (DOI 10.3847/2041-8213/ae3ddc), a candidate planetesimal-collision event around a likely young F-type star. The goal was a long-form article in which every number traces back to that source or to a second, equally named primary source.

There was no automatic release step at the end of the chain: a human read the finished text and approved it. This page describes exactly that one run, not science journalism in general, and it is not a claim that an editorial process can be automated.

The process

Five stages, from topic choice to cross-check.

  1. Topic choice. The target was a current paper with a solid numerical basis and open access. The choice fell on Tzanidakis & Davenport (2026).
  2. Full-text primary sources. The main paper was read in full through its freely accessible arXiv version, tables included. Two further sources were resolved through open metadata and preprint interfaces. No paywall was bypassed.
  3. Text. The draft was written from the sources actually read, with one traceable source per load-bearing number.
  4. Chart through the code path. The data visualization was generated from the cited raw values by script, not by an image model, and was actually rendered and looked at after every change, not just checked for syntax.
  5. Cross-check by a second agent with fresh context. A separate instance re-researched every load-bearing number on its own, without using the draft's notes, and returned its own verdict per claim.

Centerpiece

Nine interventions between draft and release.

Run log 20 July 2026

Four corrections

The research role replaced a draft figure of 11,000 light-years with the paper's value, 3,551 parsecs (about 11,583 light-years, arXiv:2603.10952); the description “Sun-like” was checked against the paper and flagged as imprecise. The star is likely a young F-type star (Teff ≈ 6,479 kelvin), not a Sun, which is a G-type star at 5,772 kelvin.

Version 1 of the data chart had a footer that collided with the axis label. The visualization role caught it at the render check, only once the actually rendered image was looked at, not in the code itself; version 2 fixed it.

Earth's Moon reference mass was stated as 7.342×10²² kilograms, attributed to “NASA fact sheets.” The NSSDC fact-sheet value is actually 7.346×10²² kilograms; 7.342 is the IAU value. The difference is 0.05 percent and has no effect on the chart, but the attribution was wrong regardless. The context-separated cross-check found this, not the author; version 3 corrected both the value and the attribution.

Two values for the emitting cross-section of the dust in the paper, 0.13 and 0.08 square AU (arXiv:2603.10952, §4.1), were described in the draft as “a reassuring agreement.” The two numbers differ by a factor of about 1.6, roughly 60 percent. The cross-check flagged the wording as flattering; the text now states the size of the gap instead of smoothing it over.

Two rejected image candidates

The AI-generated hero image first went through a self-check: six checklist items, each ticked, with an overall verdict of “CLEAN — accepted.” It was not that self-check that caught the problem, but a subsequent, separate design gate, which reviewed the same version again and found three issues, including diffraction spikes that mimicked a telescope exposure and so contradicted the image's own caption, “not observational data.” The second version drops any instrument signature; the bans moved permanently into the image template.

After switching to a newer default image-generation model, that model produced an image with more surface detail but factually wrong: a self-luminous explosion instead of an occultation, exactly the picture the article text explicitly rejects twice. The separate design gate flagged that too; the new image was discarded and the older one stayed in use.

One preserved ambiguity

The paper names two differently labeled quantities for the debris mass, a factor of roughly 12 apart: a “dust mass” of 4×10²⁰ kilograms in the abstract, unqualified, and a “total emitting mass” of 5×10²¹ kilograms in §4.1 (DOI 10.3847/2041-8213/ae3ddc), there explicitly named a conservative minimum. The cross-check confirmed that from the outside this cannot be resolved to one value; even the paper's own comparison to the Saturnian moon Enceladus sits at the low end of the range. The text carries both numbers as a range, not as one resolved value.

One claim that was stopped and later substantiated

A contextual claim about a 2014 event at a young star was plausible but had not yet been checked against a primary source; the cross-check verdict was UNVERIFIED, and the claim was not approved. Only after that was the primary source obtained: Meng et al. 2014, arXiv:1503.05609, abstract read. The claim then stood as its own sourced key claim in the second draft.

One intervention on the tool itself

A read-only audit of the image-generation tool reviewed its style presets. A mode called “schema” allowed fact-bearing diagrams to be produced by image generation instead of from code, contradicting the tool's own rule that numeric charts are rendered from data, not generated. The mode was removed.

The process facts in this section (who found what, which image was discarded, that a tool mode was removed) come from the internal run log dated 20 July 2026 and are not externally linkable. Numbers from the reviewed paper are each linked to their source.

One example, in detail

The chart that was made three times.

Below is the third and final version of the data chart. There is no before-image: reconstructing the faulty first version after the fact would itself be a third version with the layout bug artificially reintroduced, and the original layout parameter was never recorded in the log.

Data chart Version 3 of 3
Logarithmic chart: Gaia-GIC-1 debris mass, 4×10²⁰ to 5×10²¹ kilograms, compared with Enceladus, Ceres, and Earth's Moon, mass scale 10²⁰ to 10²³ kilograms.
Tap to enlarge. The labels are in English, as generated in the original run. Data: Tzanidakis & Davenport 2026, ApJL (arXiv:2603.10952), abstract & §4.1; reference masses: NASA/JPL & NSSDC fact sheets. On the range: only the upper value, 5×10²¹ kg, is described in the paper as a conservative minimum; 4×10²⁰ kg appears in the abstract without that qualifier. The “lower bound” label inside the image comes from the original run, covers both values, and was deliberately not retouched afterwards.
Generated by script from the cited raw values; no image-generation model involved.

The three versions on record, taken from the internal run log of 20 July 2026 and not externally linkable:

  1. Version 1. The footer collided with the axis label. Found by looking at the rendered image.
  2. Version 2. Layout fixed. The Moon mass was stated as 7.342×10²² kilograms and attributed to “NASA fact sheets.” The context-separated cross-check found that, not the author.
  3. Version 3. Moon mass aligned to the NSSDC value, 7.346×10²² kilograms, attribution corrected. This is the version shown.

This rendering requirement goes beyond layout bugs. In a separate, earlier run, an SVG chart fully passed an XML validity check and a geometry-overflow check and still rendered wrong, because the width of rendered text cannot be computed without an actual renderer (internal capability assessment dated 20 July 2026, a separate run).

The finding

One result for this run, not a generalization.

Nine interventions in a single run are a recorded single case, not evidence of a general pattern. In this run, the documented problems were mostly plausible inaccuracies and presentation errors. Across the eleven defined key claims: no DISPUTED verdict. Two non-load-bearing contextual claims were initially not primary-verified; one of them was later substantiated (see above). One claim is “CONFIRMED with a caveat” (the Moon attribution). The NASA NSSDC pages were not reachable during the audit; reference values were checked against standard constants instead. No hallucination finding within this audit's scope.

Four patterns carry this finding, in the vocabulary of the playbook “Orchestrating AI Agents”: a cross-check by a second agent with fresh context (“fresh context”); a reviewer that only reads and directs nothing (“read-only reviewer,” “trust boundary”); evidence that is checked before building on it, not just cited (“verify gate”); and a human who decides at the end (“human gate”). A fifth pattern carries no matching product term of its own: a chart only counts as checked after it has actually been rendered, not after a purely syntactic check.

Further reading

Where this goes next.

Broader in scope, for the whole business: The AI-First Operating Playbook →

Methodology & status

How this page is sourced.

This page describes one run, one field, one article: no claim about scaling to other topics, editorial teams, or models. Status: 20 July 2026 (run) / 21 July 2026 (this page). Numbers from the reviewed paper are checked against the primary source linked at each point. Reference masses were checked against standard constants where the NSSDC pages were unreachable. Process details come from the internal run log and are not externally linkable.

What can be checked from the outside is what is linked: the primary sources, and the numbers taken from them. The run logs stay internal, the original piece is not linked, and a first version of the chart no longer exists; this page therefore makes no before-and-after claim. It says nothing about effort, because no time was recorded during the run. This page was created with AI assistance and editorially reviewed, a transparency notice in the spirit of Art. 50 of the EU AI Regulation.