bigbounce

All worksData release

Data release · N2

Abstract

A focused, reproducible public-identifier and provenance data release for the anomaly program's historical candidate list. Deterministic join and checkpoint machinery ties each candidate to a public DESI DR1 TARGETID, enabling auditable independent follow-up. The declared 1-arcsec positional join yields 181 warning-free global-primary DESI DR1 associations: 170 at or below 0.1 arcsec and 11 lower-confidence associations between 0.1 and 1 arcsec. The sub-0.1-arcsec core is expected seed self-recovery — the cluster centroid equals the seed DESI member's own coordinates by construction, verified end-to-end rather than an independent association test. The aggregate annular shift comparison is descriptive, not a conditional false-association null or purity estimate. The release carries exact source-row provenance, explicit quality tiers, warned-row auxiliary data, checksums, and a clean-checkout validator while declining physical classification, purity, novelty, and anomaly-rate claims unsupported by the underlying historical candidate list.

Result summary

  • derived181 unique warning-free global-primary DESI DR1 TARGETIDs, partitioned exactly into 170 core and 11 lower-confidence positional associations
  • open20,299,155 eligible DESI rows → 2,468 positional parents → 2,448 global-primary rows → 181 warning-free associations
  • derivedEvery released row and all 18 carried DESI fields were re-read from the recorded FITS row and compared exactly
  • openSixteen deterministic local shifts yield 86.7 ± 14.4 parent and 76.2 ± 13.3 warning-free-primary associations within 1 arcsec; the 11-row tail is not treated as secure identity

Figures

Readiness

95%

0B / 0M / 0m / 0C open

Readiness is publication readiness only — science, evidence, review convergence, packaging, and Houston’s final sign-off. Venue and submission are tracked separately below, and never subtract from this number (directive P).

Publishing

Not part of readiness.

target venue Integrated supporting release for the rebuilt DESI anomaly flagshipstate in revision
Review evidence (193 rounds)
autolog-2026-09-04Skill/process autolog 2026-09-04: 9 commits landed via tools/skills_autolog.sh2026-09-04
anomaly-catalogue-data-release-ledger8-2026-09-03Anomaly catalogue → data release (ledger #8 answered, phase-3 v2 landed)2026-09-03
skill-improvement-2026-09-03-fable-early-commitFable/Opus science lanes must create+commit their output file within the first ~10 tool calls2026-09-03
skill-improvement-2026-09-03-no-nested-delegationSonnet execution lanes must not spawn nested background agents2026-09-03
anomaly-sample-provenance-preflight-2026-09-03Sample-provenance preflight before any GPU-billed anomaly run2026-09-03

Automated review is a gate on publication readiness, not a product. Full review timeline →

External peer review kit

Houston-driven manual round · paste prompt into any frontier LLM web UI with the PDF attached

One click to copy a referee prompt scoped to this paper. One click to download the latest PDF. Paste both into Claude / GPT-5 / Gemini / Grok / Perplexity — return findings here and the autonomous cron will close them in the next bundled hard-fix wave.

preview prompt
You are an external referee for MNRAS / Physical Review D / JCAP (target journal depends on paper).

Attached: Paper 3 v3.2.0-r17 — "Public-ID Recovery for a Historical DESI DR1 Anomaly List: 170 High-Coordinate-Consistency Core and 11 Lower-Confidence Positional Associations"
Source: pipelines/p3_anomaly_engine/paper3_apjs.tex
PDF: PDF · 17 pp · v3.2.0-r17 · updated Aug 3, 2026 · md5 477b0d83ca31f6ace3273bb19bcfcf34 · sha256 9a3769269ada4d2a5371aa447e6ce93aa55518ae2a3b13fdc3d83f0b8b779a0b — exact denominator, deduplication-order, checkpoint-digest, and viewer-binding closures included; exact r17 confirmation remains pending.

Read the FULL PDF end-to-end. Produce a referee report in MNRAS format with:

1. Recommendation: ACCEPT / MINOR REVISIONS / MAJOR REVISIONS / REJECT
2. BLOCKERS (must fix before publication) — list each with section/line + proposed fix
3. MAJORS (should fix) — same format
4. MINORS (polish) — same format
5. Strengths (>= 3 bullet points)
6. Specific scrutiny on:
   - 181 warning-free global-primary DESI DR1 TARGETIDs
   - Exact 170 core plus 11 lower-confidence positional quality-tier contract
   - Sub-0.1-arcsec core explicitly framed as expected seed self-recovery
   - Checksum-bound release validation, bounded final-hash confirmation, and ApJS metadata

CALIBRATION (do not burn findings on these known classes):
- The current date is June 2026. arXiv identifiers of the form 25xx.xxxxx and 26xx.xxxxx are VALID, already-published preprints — do not flag them as "future-dated" or "nonexistent". Verify a citation against arXiv/ADS before claiming it does not exist.
- Correction notes, retraction notices, and "an earlier version stated X" disclosures in the text are DELIBERATE transparency policy. Flag them only if their content is wrong, never for existing.
- Companion-paper citations marked "posted concurrently on arXiv" are deliberate placeholders; real arXiv IDs are inserted during the coordinated submission sequence.
- Explicitly labeled conservatism allowances, scaling estimates, ansatz/heuristic status labels, and disclosed queued follow-up computations are deliberate scoping, not oversights — flag only if the label itself is inaccurate.
- PDF text extraction can mangle math (square roots, fractions, superscripts). Before flagging "garbled" or "wrong" math, consider extraction artifacts; flag only what is visibly wrong in the rendered PDF.

VERDICT STANDARD (apply the SAME high bar a first-pass Physical Review D / MNRAS referee would — this is one of the most rigorous journals in the world):
- Assign each finding's severity (BLOCKER / MAJOR / MINOR) by your own independent referee judgment. Do NOT default to any tier, and do NOT soften a finding because the rest of the paper is strong. Do not echo this prompt's context.
- A reporting choice that headlines the more favorable of two numbers, an unstated assumption, an uncontrolled systematic, or an internal inconsistency IS a real finding — classify it honestly (MINOR at minimum), not as mere "style" or "opinion".
- Truth-audit any claim that seems off by checking it against the published .tex / on-disk artifacts before flagging (this only filters out genuine extraction artifacts — it does not lower the bar on real defects).

Reproduce this

DP3-15 held-out re-inference (structural-ceiling demonstration)cpu-only · est. 30-60 minutes (dominated by the live SPARCL re-pull of ~20,000 held-out spectra) · $0.00runnable-now
eROSITA scaler-leakage bounded control (top-298 overlap 257/298, J=0.76)unknown (930,203-source scaler-refit comparison implies at minimum cpu-heavy, plausibly gpu-24gb given the paired A/B retrain) · est. unknown until the generating script is recovered or rewritten from the committed result JSON's documented method · $0.00runnable-now
DESI 5-fold cross-validation reproducibility gate (mean pairwise Jaccard 0.862)cpu-only (5-fold BigAE proxy training over 47,000-row pool) · est. 1-3 hours (5-fold model training, 47,000-row pool) · $0.00runnable-now
Multi-survey summary / crossmatch / spatial-clustering / score-distributions (8 surveys)cpu-only · est. hours (8-survey crossmatch, external archive query throughput bound) · $0.00needs-data-restore
NANOGrav 15-yr free-spectrum PTA MCMC (real Zenodo KDE likelihood, emcee)cpu-only (32 walkers, 30 frequency bins, production run completed in 24.97s wall per committed results.json) · est. under 1 minute (committed run reports production_seconds=24.97 for 32 walkers x 10,000 production + 2,500 burn-in) · $0.00runnable-now
Planck held-out membership test + native re-inference (partial, 48/200 vs 30 expected)cpu-only for the membership test; full native re-inference requirement unknown (blocked, never re-specified post pod exit) · est. minutes for the membership test; full native re-inference is not schedulable until best_cmb_native.pt + cmb_native_patches.npy + cmb_native_all_scores.parquet are re-staged · $0.00needs-data-restore

Full reproduction manifests →

Lineage

Supporting Data Release · DESI Public-ID Recovery. This work does not claim beyond its stated target and scope above.