Review activity
A gate on readiness, not a product
Automated multi-model review is a gate on publication readiness, not a product. Rounds stop when the remaining findings are genre or venue (directives R2, P). Raw machine events (dispatches, closures) stream at /activity — this page is the curated review-loop story.
Verdict grid · newest round left
External referee verdicts
Active legs only (directive M-AMENDED): Grok API + Gemini API, plotted against the historical six-paper board. The ChatGPT column is frozen while directive N's Codex/OpenAI pause stands — shown dimmed, never deleted or faked.
Automated-review diagnostic only. Per directive M-AMENDED (2026-07-23) this counts ACTIVE legs only: Grok + Gemini (grid columns) and the Claude INT leg (verdicts in round notes) — 3 legs × 6 papers. The GPT column is excluded while paused (directive N, since 2026-07-16); its history stays displayed. An ACCEPT here is not journal acceptance, and 100% is not required to submit a paper.
| Work | CONFIRM-2026-07-22 | M45 | M42 | M27 | M25+M26 | M24 | M23 | M18 | M3 | M2 |
|---|---|---|---|---|---|---|---|---|---|---|
| P1A | —Am | ——— | Rm— | ——— | ——— | ——— | RM— | Rm— | ——— | ——— |
| P1B | —mm | ——— | ——— | ——— | ——— | ——— | ——— | ——— | ——— | ——— |
| P2 | —Mm | Mm— | ——— | ——— | Rm— | ——— | RM— | Rm— | ——— | ——— |
| P3 | —AM | ——— | RM— | —M— | ——— | RM— | ——— | ——— | R—— | —M— |
| P4 | —mm | ——— | Rm— | ——— | —m— | ——— | ——— | ——— | ——— | Rm— |
| P5 | —Am | Mm— | ——— | ——— | Mm— | ——— | ——— | ——— | —m— | Mm— |
ChatGPT column is frozen, not counted — paused under standing directive N; history is preserved, never deleted or faked.
Historical board versions/caps: P1A 95, P1B 95, P2 95, P3 95, P4 95, P5 95. The live-lineup works (A3, P4′, P1N) are not yet columns in this historical grid — their round-by-round evidence is in the timeline below and their readiness is on /status.
Publication status
What's left before publication
What's left before publication
0/6 signed off
0 waiting on Houston · 6 on the agents
evidence as of Sep 7, 2026
How to read this. Publication readiness is science closure + evidence & reproducibility + automated review convergence + packaging, and then Houston's own final read — the last 5%. A paper marked Houston needs no further math, compute, GPU/CPU runs or new data; a paper marked agent has one named item still owned by the loop. The trailing stamp shows which exact PDF the newest automated review board actually read — ✓ means the current one, ↗ means an earlier one.
Publishing is a separate phase. arXiv endorsement, venue choice, submission clicks and journal / independent human review come after 100% and never subtract from readiness. See the publishing checklist →
Gap and skills
The review machinery, self-improving
Substantive findings only the external tier caught, and the pattern/prompt-rule catalog those findings get mined into.
Round timeline · newest first · showing 60
Every round, truth-audit, closure, and skill upgrade
One line per event: date, kind, what changed, receipt link. Skill-improvement entries carry a quiet marker.
Full history (append-only, 448 rounds total) in reviewTimeline.ts.