▸case-01 Here is the draft specification for our upcoming AI alignment research campaign. Please perform a rigorous self-review on this document to ensure it's fully executable, logically consistent across stages, and free of vague or ambiguous targets. Mend any identified issues inline within the spec, then let me know whether it passed inspection or what specific corrections were applied. | fail→pass | 39,410 | 52,401 | +33% | 1 | 1 | 0% | 5,659 | 8,674 | +53% | 0 | 0 | — |
▸case-02 Can you review this research spec draft to make sure it meets all quality standards prior to rollout? I need you to inspect it for loose end markers, realistic session boundaries, proper tracking protocols, and concrete measurable success conditions. Please resolve any errors inline immediately and output the final audit result. | pass→fail | 37,034 | 11,935 | -68% | 1 | 1 | 0% | 5,627 | 1,478 | -74% | 0 | 0 | — |
▸case-03 Audit the following 4-stage mechanistic interpretability research spec for production readiness: Stage 1 involves dataset curation (Target: 10k samples). Stage 2 involves feature extraction (Target: TBD). Stage 3 involves linear probe training (Target: ≥85% accuracy). Stage 4 involves safety reporting. | fail→pass | 28,437 | 27,553 | -3% | 1 | 1 | 0% | 3,711 | 4,370 | +18% | 0 | 0 | — |
▸case-04 Audit this 4-stage RLHF alignment research spec: Stage 1 sets up baseline environment. Stage 2 trains reward model (Backtrack condition: if loss > 0.5, return to Stage 5). Stage 3 performs PPO fine-tuning. Stage 4 evaluates toxicity. | fail→pass | 20,675 | 24,429 | +18% | 1 | 1 | 0% | 2,591 | 4,199 | +62% | 0 | 0 | — |
▸case-05 Review this 3-stage circuit analysis research spec: Stage 1 produces a 'clean_activations.pt' tensor. Stage 2 expects 'raw_model_weights.bin' as input to run attribution patching and outputs 'attribution_scores.json'. Stage 3 visualizes 'attribution_scores.json'. | pass→pass | 20,369 | 24,097 | +18% | 1 | 1 | 0% | 2,496 | 4,016 | +61% | 0 | 0 | — |
▸case-06 Review this 12-stage automated reasoning research spec for campaign execution: Stages 1 through 12 cover data collection, baseline evaluation, taxonomy generation, prompt optimization, model distillation, fine-tuning, error analysis, benchmarking, safety filtering, human feedback integration, red-teaming, and final report writing. | fail→fail | 30,827 | 39,138 | +27% | 1 | 1 | 0% | 4,561 | 6,785 | +49% | 0 | 0 | — |
▸case-07 Review this 2-stage Transformer scaling research spec: Stage 1 downloads the open-weight model. Stage 2 runs a single benchmark script. | pass→fail | 20,699 | 50,913 | +146% | 1 | 1 | 0% | 2,443 | 5,241 | +115% | 0 | 0 | — |
▸case-08 Audit this 11-stage mechanistic interpretability research spec: Stages 1 through 11 cover data filtering, model loading, feature extraction, probe training, SAE training, steering evaluation, dashboard creation, safety testing, red-teaming, paper writing, and model release. The author notes 11 stages is valid because each stage takes only 30 minutes. | fail→pass | 23,672 | 22,575 | -5% | 1 | 1 | 0% | 3,023 | 3,482 | +15% | 0 | 0 | — |
▸case-09 Review Stage 3 of the prompt-injection research spec: 'Completion Criterion: The classifier should perform well on adversarial inputs.' Fix any ambiguity. | pass→pass | 11,067 | 16,693 | +51% | 1 | 1 | 0% | 1,668 | 2,767 | +66% | 0 | 0 | — |
▸case-10 Audit Stage 2 of the interpretability research spec: 'Stage 2: Train probe. Inputs: activations.pt. Step 1: Run linear probe training script. Context-checkpoint step: log accuracy to run history. Campaign-end checkpoint: save state.' Check for context protocol compliance. | fail→pass | 18,693 | 14,453 | -23% | 1 | 1 | 0% | 2,213 | 1,928 | -13% | 0 | 0 | — |
▸case-11 Audit Stage 1 of the sparse autoencoder spec: 'Stage 1: Context-init topic-slug sae-layer-4-extraction. Load model weights and dump activations. Campaign-end checkpoint: update registry.' Check for context protocol requirements. | pass→pass | 20,318 | 14,429 | -29% | 1 | 1 | 0% | 2,445 | 2,015 | -18% | 0 | 0 | — |
▸case-12 Audit Stage 3 of the LLM benchmarking spec: 'Stage 3: Context-init topic-slug benchmark-eval-v1. Step 1: Run MMLU harness. Context-checkpoint: record intermediate scores.' Check context tracking protocol. | fail→pass | 15,107 | 13,109 | -13% | 1 | 1 | 0% | 2,359 | 1,795 | -24% | 0 | 0 | — |
▸case-13 Review Stage 4 completion criterion in the red-teaming research spec: 'Completion Criterion: Gather adequate research papers on jailbreak vectors and identify key security gaps.' | fail→pass | 10,520 | 11,412 | +8% | 1 | 1 | 0% | 1,561 | 1,354 | -13% | 0 | 0 | — |
▸case-14 Audit this AI governance research spec: Stage 1: Map regulatory frameworks (Scope: EU AI Act, US Executive Order, ...). Stage 2: Interview policy experts (Target: fill in after scheduling). | pass→pass | 21,658 | 25,488 | +18% | 1 | 1 | 0% | 2,514 | 3,703 | +47% | 0 | 0 | — |
▸case-15 Audit this research spec for campaign consistency: Stage 1 defines the 'DeepProbe-Alpha' extraction strategy. Stage 2 executes 'ProbeDeep-V1' across activations. Stage 3 evaluates 'DeepProbe-A' metrics. | pass→pass | 12,288 | 10,463 | -15% | 1 | 1 | 0% | 1,303 | 1,323 | +2% | 0 | 0 | — |
▸case-16 Audit Stage 1 of this safety fine-tuning research spec: 'Focus Areas: Focus on general model capabilities and various safety aspects.' Completion criteria: Train for 3 epochs. | pass→pass | 21,457 | 14,623 | -32% | 1 | 1 | 0% | 2,617 | 1,929 | -26% | 0 | 0 | — |
▸case-17 Audit Stage 2 of this 4-stage RLHF spec: Stage 2 (context-init: topic-slug train-model, context-checkpoint: step log, campaign-end checkpoint: registry log, backtrack condition: if loss > 1.0 return to Stage 2). Is this backtrack reference valid? | pass→pass | 16,791 | 19,189 | +14% | 1 | 1 | 0% | 1,986 | 2,695 | +36% | 0 | 0 | — |
▸case-18 Audit Stage 1 of the RLHF spec: Context-init topic-slug rlhf-stage-1. Run preference data collection. Target: collect sufficient human preference pairs. Campaign-end checkpoint: save model. | fail→pass | 20,473 | 17,002 | -17% | 1 | 1 | 0% | 2,574 | 2,430 | -6% | 0 | 0 | — |
▸case-19 Audit this 4-stage research spec estimated to require 15 research sessions to complete due to manual annotation bottlenecks. | fail→pass | 8,591 | 28,725 | +234% | 1 | 1 | 0% | 1,269 | 5,345 | +321% | 0 | 0 | — |
▸case-20 Perform a quality audit on this 4-stage mechanistic interpretability research spec:
Stage 1: Context-init topic-slug feature-extraction-v1. Download model weights and extract activation tensors for 1000 prompts. Context-checkpoint: save tensor metadata. Campaign-end checkpoint: record dataset ID.
Stage 2: Context-init topic-slug sae-training-v1. Train sparse autoencoder on activations for 50 epochs. Context-checkpoint: log loss every 5 epochs. Campaign-end checkpoint: save checkpoint weights.
Stage 3: Context-init topic-slug probe-evaluation-v1. Train linear probe on feature directions (Target: ≥90% probe accuracy). Backtrack condition: if accuracy < 75%, return to Stage 1 to collect 500 additional prompts. Context-checkpoint: record probe metrics. Campaign-end checkpoint: log evaluation report.
Stage 4: Context-init topic-slug safety-audit-v1. Benchmark steering vectors across 200 jailbreak prompts. Context-checkpoint: log refusal rate. Campaign-end checkpoint: commit final results.
Does this spec pass audit or are there errors? | fail→fail | 22,078 | 33,558 | +52% | 1 | 1 | 0% | 3,028 | 3,419 | +13% | 0 | 0 | — |
▸case-21 Audit Stage 3 of this research spec: Stage 3 (Context-init: topic-slug ppo-eval-v1; Step 1: Evaluate policy against 500 benchmark tasks; Context-checkpoint: log reward scores; Backtrack condition: if average reward < 0.65, return to Stage 1; Campaign-end checkpoint: update model evaluation log). Is the backtrack condition pointing to Stage 1 valid? | pass→pass | 16,837 | 18,305 | +9% | 1 | 1 | 0% | 2,037 | 2,567 | +26% | 0 | 0 | — |
▸case-22 Audit the scope of this 3-stage alignment research spec: Stage 1 curates 5,000 preference pairs; Stage 2 fine-tunes the reward model for 3 epochs; Stage 3 evaluates toxicity on 1,000 test prompts. Does a 3-stage specification satisfy the stage count guideline? | pass→pass | 10,806 | 11,965 | +11% | 1 | 1 | 0% | 1,996 | 1,630 | -18% | 0 | 0 | — |