▸case-01 Here is a passage from our paper's peer-review rebuttal draft:
'> **Reviewer's Comment**: Baseline accuracies seem lower than expected.
**Response**: We appreciate the reviewer's observation. It is important to note that the baseline results we reported in Table 2 were obtained using the standard public benchmark settings provided by the original authors without any fine-tuning.'
This draft response is too wordy. Please make it significantly more concise and provide your output as a JSON object with a 'text' property. | fail→pass | 2,764 | 3,244 | +17% | 1 | 1 | 0% | 647 | 1,092 | +69% | 0 | 0 | — |
▸case-02 We need to trim our peer-review rebuttal draft for length. The text contains a template placeholder token: 'Dear {{reviewerId}}, we acknowledge that the computational overhead in epoch 3 is high, but as shown in Table 1, our method improves latency by 14%.' Please condense this response to be more concise and output a JSON object with a 'text' key. | pass→pass | 3,242 | 3,685 | +14% | 1 | 1 | 0% | 716 | 1,193 | +67% | 0 | 0 | — |
▸case-03 Please make the following rebuttal paragraph concise: 'As established by Smith et al. (2022), the loss function $L_{total} = L_{cls} + \lambda L_{reg}$ optimizes both classification accuracy and bounding box regression simultaneously, which directly prevents over-smoothing during late training phases.' Output JSON with a 'text' key. | pass→pass | 4,815 | 5,086 | +6% | 1 | 1 | 0% | 820 | 1,231 | +50% | 0 | 0 | — |
▸case-04 Please shorten this draft response while keeping it accurate: 'In our ablation experiment with $N=500$ samples, we observed a peak memory consumption of 12.4 GB and an average throughput of 450 tokens/sec across 4 NVIDIA A100 GPUs.' Output strict JSON with a 'text' field. | pass→pass | 5,533 | 4,115 | -26% | 1 | 1 | 0% | 1,180 | 1,278 | +8% | 0 | 0 | — |
▸case-05 The author response currently says: 'We did not include runtime benchmarks on mobile devices because our current implementation targets server-class hardware.' Condense this sentence, but do not promise future experiments or additional code releases that were not in the original text. Return a JSON object with a 'text' property. | pass→pass | 5,844 | 4,180 | -28% | 1 | 1 | 0% | 945 | 992 | +5% | 0 | 0 | — |
▸case-06 Please condense this rebuttal segment without softening or changing our stance: 'We strongly disagree with the assertion that our dataset lacks diversity, because our collection spans 12 distinct geographic regions and 5 separate climate categories.' Format the response as a JSON object with a 'text' key. | pass→pass | 6,548 | 3,499 | -47% | 1 | 1 | 0% | 1,012 | 1,045 | +3% | 0 | 0 | — |
▸case-07 This sentence in our paper rebuttal is already succinct: 'Table 3 shows our model outperforms ResNet-50 by 2.3% top-1 accuracy.' Condense this text into JSON with a 'text' field, ensuring minimal edits are made since it is already concise. | pass→pass | 6,354 | 2,973 | -53% | 1 | 1 | 0% | 1,323 | 937 | -29% | 0 | 0 | — |
▸case-08 Condense the following bulleted list in our rebuttal while preserving its markdown list structure:
- First, we conducted a 5-fold cross-validation experiment to verify model stability.
- Second, we computed $p$-values ($p < 0.001$) across all split folds to confirm statistical significance.
Return the output in a JSON object with a 'text' key. | pass→pass | 3,183 | 5,565 | +75% | 1 | 1 | 0% | 638 | 1,299 | +104% | 0 | 0 | — |
▸case-09 Condense this wordy rebuttal sentence: 'It goes without saying that we feel it is arguably possible that our hyperparameter choices might potentially have influenced the final validation accuracy to a small degree.' Return a JSON object with a 'text' key. | pass→pass | 3,600 | 2,859 | -21% | 1 | 1 | 0% | 555 | 874 | +57% | 0 | 0 | — |
▸case-10 Shorten this statement from our paper rebuttal: 'We evaluated our pipeline using the Stanford Question Answering Dataset (SQuAD v2.0) and compared results against BERT-Large and RoBERTa-Large.' Output JSON with a 'text' property. | pass→pass | 3,888 | 3,719 | -4% | 1 | 1 | 0% | 663 | 1,034 | +56% | 0 | 0 | — |
▸case-11 Please condense the following response text while keeping reviewer quote blocks intact:
> **Reviewer Comment**: 'The latency numbers in Section 4.2 require clarification.'
**Response**: In Section 4.2, latency was measured on an RTX 3090 GPU over 1000 trials, yielding a mean latency of 14.2 ms.
Provide the result as a JSON object with a 'text' key. | pass→pass | 3,698 | 6,215 | +68% | 1 | 1 | 0% | 872 | 1,549 | +78% | 0 | 0 | — |
▸case-12 Please condense this draft: 'While our preliminary findings suggest a potential performance gain on small datasets, we cannot definitively conclude that this trend generalizes to large-scale data without further empirical testing.' Output JSON with a 'text' property. | pass→pass | 3,223 | 2,858 | -11% | 1 | 1 | 0% | 604 | 838 | +39% | 0 | 0 | — |
▸case-13 Condense this repetitive rebuttal passage: 'We chose AdamW because it provided superior stability during initial training runs. Furthermore, AdamW was selected due to its demonstrated stability compared to SGD in early epochs.' Return JSON with a 'text' key. | pass→pass | 2,652 | 2,594 | -2% | 1 | 1 | 0% | 523 | 842 | +61% | 0 | 0 | — |
▸case-14 Please shorten this draft response: 'As defined in Equation 3, where $\theta_{new} = \theta_{old} - \alpha \nabla L$, the learning rate parameter \alpha was fixed at $10^{-4}$ throughout all fine-tuning runs.' Output JSON with a 'text' field. | pass→pass | 4,096 | 4,063 | -1% | 1 | 1 | 0% | 673 | 1,256 | +87% | 0 | 0 | — |
▸case-15 Condense this two-paragraph rebuttal while preserving the two-paragraph layout:
Paragraph 1: Regarding the baseline dataset, we used the standard 80-20 train-test split established in prior benchmark studies.
Paragraph 2: Regarding hardware specifications, all experiments were conducted on 8 NVIDIA V100 GPUs with 32 GB VRAM each.
Output JSON with a 'text' key. | pass→pass | 5,855 | 4,798 | -18% | 1 | 1 | 0% | 1,291 | 1,210 | -6% | 0 | 0 | — |
▸case-16 Condense the following paper rebuttal text: 'In summary, our contribution is a novel attention mechanism that reduces computational complexity from $O(N^2)$ to $O(N \log N)$ while maintaining identical top-1 accuracy on ImageNet.' Format the output as strict JSON with a 'text' key. | pass→pass | 4,344 | 3,616 | -17% | 1 | 1 | 0% | 896 | 1,177 | +31% | 0 | 0 | — |
▸case-17 Shorten this rebuttal sentence: 'Our training strategy follows standard practices in domain adaptation literature.' Do not insert specific paper citations or author names that were not present in the original sentence. Output JSON with a 'text' key. | pass→pass | 3,343 | 2,945 | -12% | 1 | 1 | 0% | 630 | 788 | +25% | 0 | 0 | — |
▸case-18 Condense this rebuttal text: 'We have updated Section 3.2, Figure 4, and Appendix B to include the additional ablation results requested by the reviewer.' Output JSON with a 'text' key. | pass→pass | 3,070 | 2,885 | -6% | 1 | 1 | 0% | 519 | 869 | +67% | 0 | 0 | — |
▸case-19 We need to expand our rebuttal response to add context and detailed explanations. Here is our brief draft: 'We used AdamW with learning rate 1e-4.' Please write a detailed, expanded 3-paragraph explanation justifying this hyperparameter choice, citing common optimization theories in deep learning. | pass→fail | 12,553 | 5,287 | -58% | 1 | 1 | 0% | 2,003 | 1,229 | -39% | 0 | 0 | — |
▸case-20 Reviewer 2 asked: 'Why did you not compare against Method X?' We have not drafted any response yet. Please write a brand new response draft addressing why Method X was omitted due to code unavailability. | pass→pass | 9,721 | 5,703 | -41% | 1 | 1 | 0% | 1,543 | 1,223 | -21% | 0 | 0 | — |
▸case-21 Analyze the sentiment and tone of the following reviewer comment: 'The empirical results are convincing, but the mathematical formulation in Section 3 is overly obscure and lacks intuition.' Provide a breakdown of positive vs negative feedback points. | pass→fail | 8,249 | 8,110 | -2% | 1 | 1 | 0% | 1,350 | 1,704 | +26% | 0 | 0 | — |
▸case-22 Please make this rebuttal text concise: 'Our method uses Convolutional Neural Networks (CNN) combined with Long Short-Term Memory (LSTM) networks to process multimodal time-series telemetry data.' Output JSON with a 'text' property. | fail→pass | 2,779 | 6,005 | +116% | 1 | 1 | 0% | 523 | 1,435 | +174% | 0 | 0 | — |