The scientist retains its previous code, results, manuscripts, and feedback so that revisions accumulate and the reviewer retains no history, preventing earlier judgments or scores from biasing its assessment of the current version.
Given a rejected paper, its source repository, and current feedback, the scientist is instructed to address every reviewer concern while determining how each should be resolved. Depending on the feedback, this may require inspecting the repository and manuscript, searching the literature, editing implementation, reproducing results, adding baselines and presenting ablations, executing new experiments and analyzing the results and updating the manuscript.
The central research contribution is treated as fixed; if addressing a concern would require changing that contribution, the scientist records it as unresolved rather than reframing the work as a different project.
Revision conditions
Human-initializedOriginal written venue reviews guide the first round; a fresh AI review guides each later revision.
AI-initializedA fresh, independent AI review supplies the initial feedback and guides each later revision.
Autonomous self-revisionThe scientist receives the same fixed self-review prompt in every round.
- Evaluation set
- 25 rejected and 5 borderline-accepted ICLR papers
- Revision rounds
- V₀ original; V₁–V₅ successive revisions
- Compute
- Approximately 9 hours per paper for the complete five-round revision run on a 96 GB VRAM GPU
- External evaluation
- Stanford Reviewer scores the human-initialized trajectory only