AppliedScientistRead paper

AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing

Vidushee VatsKarun SharmaShengzhi LiShichao Pei

Can this process of scientific review and revision itself be automated?

We present AppliedScientist, a closed-loop system that couples an autonomous AI scientist with an AI reviewer, and evaluate it by iteratively revising rejected papers from a range of research subfields.

AI Reviewer-guided revision on 30 ICLR papers

Mean score increase assigned by our AI Reviewer

+1.16score points
5.37 original6.53 best revision

Across 30 ICLR papers, reviewer-guided revision consistently improved papers more than autonomous self-revision with a fixed prompt.

Independent evaluation

+0.65score points

Stanford Reviewer

5.50 original → 6.15 best revision

Resolved through revision

85.3%

execution-related weaknesses resolved

128 of 150 across the evaluation set

01 Methodology

AppliedScientist consists of two components

A scientist that revises the implementation, experiments, and manuscript, and an AI reviewer that independently reviews each revision and provides feedback for the next iteration.

Supporting evidence

Experimental setup

The scientist retains its previous code, results, manuscripts, and feedback so that revisions accumulate and the reviewer retains no history, preventing earlier judgments or scores from biasing its assessment of the current version.

Given a rejected paper, its source repository, and current feedback, the scientist is instructed to address every reviewer concern while determining how each should be resolved. Depending on the feedback, this may require inspecting the repository and manuscript, searching the literature, editing implementation, reproducing results, adding baselines and presenting ablations, executing new experiments and analyzing the results and updating the manuscript.

The central research contribution is treated as fixed; if addressing a concern would require changing that contribution, the scientist records it as unresolved rather than reframing the work as a different project.

Revision conditions

Human-initialized

Original written venue reviews guide the first round; a fresh AI review guides each later revision.

AI-initialized

A fresh, independent AI review supplies the initial feedback and guides each later revision.

Autonomous self-revision

The scientist receives the same fixed self-review prompt in every round.

Evaluation set
25 rejected and 5 borderline-accepted ICLR papers
Revision rounds
V₀ original; V₁–V₅ successive revisions
Compute
Approximately 9 hours per paper for the complete five-round revision run on a 96 GB VRAM GPU
External evaluation
Stanford Reviewer scores the human-initialized trajectory only
Figure 1. AppliedScientist revision loop and phase-structured AI Reviewer.
Figure 1. The scientist iteratively verifies reviewer suggestions, runs experiments, analyzes results, and updates the manuscript. The review component independently reads each revision, searches the literature, assesses its technical quality and significance, and returns structured feedback for the next round.
02Reviewer evaluation

Reviewer validity is a prerequisite for interpreting downstream improvements

The reviewer has two roles: it guides the next revision and scores its outcome. We therefore evaluate the reviewer before using its scores to assess whether papers improve.

65%70%75%0.480.520.560.60Spearman ρ with mean human ratingAccept/reject agreementour reviewer (DeepSeek V4 Flash)our reviewer (Kimi K3)Stanford Reviewerour reviewer (MiniMax M2.7)

our reviewer, run on three backbone models Stanford Reviewer

Score alignment with venue reviews on a set of 650 ICLR 2020–26 papers.

What the comparison shows

Our Reviewer with DeepSeek V4 Flash shows the strongest alignment with venue reviews on both measures.

The Stanford Reviewer falls between the strongest and weakest backbones. Systems farther toward the upper right align more closely with venue ratings and decisions.

Supporting evidence

Reviewer evaluation

Table 1. Pairwise preference evaluation of our AI Reviewer against existing automated reviewers on 650 ICLR papers from 2020–26. A search-enabled LLM judge compares anonymized review pairs using the evaluation rubric of Xu et al. Each cell gives the percentage of comparisons won by our reviewer and lost to the baseline; the remainder are ties.
BaselineTechnical accuracyConstructive valueAnalytical depthSignificance assessment
Fine-tuned
CycleReviewer-8B100 / 0100 / 099 / 1100 / 0
CycleReviewer-70B99.2 / 0.4100 / 0100 / 099.6 / 0.2
DeepReviewer-7B96.5 / 2.199.4 / 0.499.2 / 0.697.8 / 1.1
DeepReviewer-14B91.2 / 3.796 / 2.294.7 / 2.697.9 / 1.3
LLM
Gemini 3.1 Flash60.6 / 21.588.2 / 6.675.1 / 9.974.2 / 4.6
Gemini 3.1 Pro62.2 / 11.783.9 / 6.571.6 / 970.7 / 5.1
DeepSeek V4 Flash74.3 / 10.992.5 / 7.178.9 / 7.482.3 / 3.9
DeepSeek V4 Pro78.6 / 7.596.5 / 379.2 / 5.180.4 / 3.2
LLM with search
Gemini 3.1 Flash55.8 / 12.885.7 / 8.570.5 / 10.272.8 / 6.3
Gemini 3.1 Pro56.3 / 14.680.1 / 7.668.4 / 8.866.9 / 5.8
DeepSeek V4 Flash71.4 / 15.389.4 / 776.9 / 7.580.1 / 5
DeepSeek V4 Pro72.9 / 17.693.1 / 6.573.8 / 8.179.7 / 6.4
Multi-agent with search
Agent Review (Gemini 3.1 Pro)52.1 / 18.989.5 / 10.460.7 / 10.865.6 / 7.9
Agent Review (DeepSeek V4 Flash)61.7 / 20.592.3 / 472.8 / 11.668.5 / 8.7
AI Scientist v2 (Gemini 3.1 Pro)50.2 / 22.384.5 / 12.158.5 / 12.562.2 / 9.5
AI Scientist v2 (DeepSeek V4 Flash)56.7 / 24.191.2 / 7.288.3 / 8.667.1 / 10.3
DeepReviewer-v2 (StepFun 3.5 Flash)63.4 / 16.894.6 / 2.591.8 / 5.477.3 / 11.6
Stanford Agent Reviewer48.5 / 24.844.3 / 45.752.7 / 1454.6 / 10.8
03Results and discussion

Does Reviewer Guidance Improve Papers?

Reviewer scores improve steadily across successive rounds

In the human-initialized condition, the original venue reviews guide V₁ and fresh feedback from our reviewer guides each later revision. Autonomous self-revision receives the same fixed prompt in every round.

3.54.04.55.05.56.06.5V0V1V2V3V4V5Revision roundMean score6.17human-initialized4.38self-review prompt5.92Stanford (evaluation only)
Mean score by saved manuscript version. V₀ is the original manuscript; V₁–V₅ are successive revisions.
04Results and discussion

What Can Revision Improve?

What kinds of reviewer criticisms can revision actually resolve?

AppliedScientist addresses 128 of 150 execution weaknesses, but resolves only 2 of 18 idea weaknesses.

Resolution rate of weaknesses identified in the original venue reviews. Intervals show Wilson 95% confidence intervals. Select a category to inspect an example.

Example · Under the Influence

Can the reader follow the claim?

Unclear definitions, notation, organization, or interpretation of results.

Original venue review
The metric definitions were difficult to follow, and the connection between the definitions and Table 1 was unclear.
What AppliedScientist did
  • Worked interpretation of Table 1
  • Undefined GPT-5 case explained
  • Metric edge cases specified
Revised manuscript, p. 6 · worked interpretation of Table 1 and the undefined GPT-5 case.
Crop from page 6 of the revised Under the Influence manuscript showing the added interpretation of Table 1 and missing-data explanation
Under the InfluenceRevised manuscript, p. 6 · worked interpretation of Table 1 and the undefined GPT-5 case.
How “resolved” was determined

Each weakness from the original venue reviews was checked against the saved execution trajectory and revised manuscripts. This matching is separate from the AI Reviewer, which reviews each version without access to earlier reviews.

Supporting evidence

Rejection reasons and novelty verification

We randomly sample 500 rejected ICLR papers and use Gemini 3.1 Pro to classify every reviewer criticism as either an execution issue or an idea issue.

48%execution35%idea17%both
05Revision trajectories and outputs

AppliedScientist outputs

Revision trajectories and annotated manuscripts

The annotations highlight the changes made by AppliedScientist during revision.

610recorded eventsSee how one paper changed across five revision roundsBegin with the four human reviews of the original submission, then follow the reviewer feedback, experiments, and manuscript produced in each round.Open revision replay
  1. Annotated page 1 from Context is the Key: Backdoor Attacks for In-Context Learning with Vision TransformersAnnotated page 101Headline paperContext is the Key: Backdoor Attacks for In-Context Learning with Vision TransformersOpen full annotated PDF
  2. Annotated page 1 from Under the Influence: Quantifying Persuasion and Vigilance in Large Language ModelsAnnotated page 102Persuasion and vigilanceUnder the Influence: Quantifying Persuasion and Vigilance in Large Language ModelsOpen full annotated PDF
  3. Annotated page 3 from Memorization or Interpolation? What Perturbation Sensitivity Actually Detects in Language ModelsAnnotated page 303Memorization and interpolationMemorization or Interpolation? What Perturbation Sensitivity Actually Detects in Language ModelsOpen full annotated PDF