Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing or auditing the evaluation of an ACM CoNEXT paper — matching evidence to claim shape with real testbeds and deployments, honest and tuned baselines, measurement statistics and uncertainty, trace and config provenance, and contamination-aware ablations for ML-for-networking work.
.claude/skills/brycewang-stanford-conext-experiments/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -31% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 39% | 0% |
Build the evaluation a networking reviewer will actually interrogate. CoNEXT's evidence culture is systems-and-measurement: claims are backed on the real target platform — a testbed, deployment, or trace — with honest baselines and reported uncertainty, not simulation standing in for hardware or a single number with no variance. Because a one-shot major revision is decided on a list of minimum necessary changes, an evaluation gap you leave now often becomes a mandatory fix under a tight window later.
| Claim shape | Evidence CoNEXT expects | |---|---| | A mechanism is faster/cheaper on real hardware | A run on the real target (switch, NIC, kernel, testbed) under its real constraints, vs. a tuned baseline, with effect sizes | | A phenomenon exists in the wild | A measurement campaign with documented vantage points, capture dates, and a reproducible extraction methodology | | An architecture scales | Scalability evidence (real deployment or faithful emulation) across the relevant range, not a point claim | | An operator intervention helps | Evidence at operationally relevant scale, with the counterfactual measured or bounded | | A learned component adds value | An ablation isolating the learned part from the mechanism, plus a contamination check |
the switch (or a faithful hardware testbed); if about a protocol on real paths, use a testbed or trace-driven replay. Simulation-only systems claims are a classic CoNEXT revision risk.
link rates, buffer depths, and background traffic. A reviewer who cannot picture the setup cannot trust the numbers.
not a strawman default. "The baseline is not tuned/fair" is one of the most common CoNEXT push-backs.
not just means). Tail behavior is often the point in networking.
uncorrected multiple testing.
distribution pre-empts questions a mean hides.
exact configs used. These cannot be reconstructed after the fact — pin them at collection time.
be released (this interacts with double-anonymity and reproducibility).
If the paper uses a learner or LLM on networking data:
reviewer sees its marginal value over the mechanism.
training, and that a temporal split reflects real deployment.
calls re-samples rather than reproduces (see conext-reproducibility).
paper may belong at an ML venue (see conext-topic-selection).
text[Claim coverage] every claim has a matching measurement on the real target? yes/no [Platform realism] real hardware/testbed/trace, or simulation standing in? note each [Baselines] strongest reasonable, equally tuned, config reported? yes/no [Uncertainty] multiple runs, CIs/CDFs, corrected comparisons? yes/no [Provenance] vantage points, dates, configs, firmware/OS pinned? yes/no [ML checks] ablation + contamination guard + model-swap survives? n/a or yes/no
text[Evaluation status] solid / gaps [Claim-evidence matrix] <claim -> measurement + platform + baseline + uncertainty> [Platform] real target used? emulation justified? [Provenance] traces/configs/firmware pinned for reproducibility [Revision risk] <the gap most likely to become a minimum-necessary change>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 30,354 | 18,293 | -40% | 1 | 1 | 0% | 4,215 | 3,375 | -20% | 0 | 0 | — |
case-02 | fail→pass | 25,090 | 24,845 | -1% | 1 | 1 | 0% | 3,591 | 3,630 | +1% | 0 | 0 | — |
case-03 | fail→fail | 36,549 | 22,332 | -39% | 1 | 1 | 0% | 5,141 | 3,802 | -26% | 0 | 0 | — |
case-04 | pass→pass | 17,282 | 11,528 | -33% | 1 | 1 | 0% | 2,198 | 3,328 | +51% | 0 | 0 | — |
case-05 | pass→pass | 27,036 | 17,436 | -36% | 1 | 1 | 0% | 3,213 | 3,142 | -2% | 0 | 0 | — |
case-06 | pass→fail | 22,006 | 22,175 | +1% | 1 | 1 | 0% | 2,871 | 3,998 | +39% | 0 | 0 | — |
case-07 | fail→pass | 37,967 | 22,011 | -42% | 1 | 1 | 0% | 5,086 | 3,782 | -26% | 0 | 0 | — |
case-08 | pass→pass | 18,094 | 15,765 | -13% | 1 | 1 | 0% | 2,934 | 2,788 | -5% | 0 | 0 | — |
case-09 | pass→pass | 41,753 | 16,290 | -61% | 1 | 1 | 0% | 5,452 | 2,825 | -48% | 0 | 0 | — |
case-10 | fail→fail | 43,060 | 20,474 | -52% | 1 | 1 | 0% | 5,757 | 3,269 | -43% | 0 | 0 | — |
case-11 | pass→pass | 21,852 | 18,961 | -13% | 1 | 1 | 0% | 3,397 | 3,306 | -3% | 0 | 0 | — |
case-12 | pass→pass | 31,380 | 23,544 | -25% | 1 | 1 | 0% | 4,051 | 2,899 | -28% | 0 | 0 | — |
case-13 | pass→pass | 24,334 | 16,937 | -30% | 1 | 1 | 0% | 3,004 | 3,064 | +2% | 0 | 0 | — |
case-14 | pass→pass | 32,324 | 15,951 | -51% | 1 | 1 | 0% | 4,394 | 2,849 | -35% | 0 | 0 | — |
case-15 | fail→pass | 27,799 | 24,160 | -13% | 1 | 1 | 0% | 3,216 | 3,557 | +11% | 0 | 0 | — |
case-16 | pass→pass | 26,144 | 11,744 | -55% | 1 | 1 | 0% | 2,730 | 2,722 | -0% | 0 | 0 | — |
case-17 | fail→fail | 27,010 | 21,970 | -19% | 1 | 1 | 0% | 2,846 | 3,954 | +39% | 0 | 0 | — |
case-18 | pass→pass | 26,290 | 21,593 | -18% | 1 | 1 | 0% | 4,012 | 3,221 | -20% | 0 | 0 | — |
case-19 | fail→fail | 50,936 | 22,423 | -56% | 1 | 1 | 0% | 5,829 | 3,243 | -44% | 0 | 0 | — |
case-20 | pass→pass | 30,156 | 23,684 | -21% | 1 | 1 | 0% | 4,424 | 3,500 | -21% | 0 | 0 | — |
case-21 | pass→pass | 34,437 | 16,461 | -52% | 1 | 1 | 0% | 4,292 | 2,599 | -39% | 0 | 0 | — |
case-22 | fail→pass | 35,707 | 21,489 | -40% | 1 | 1 | 0% | 4,984 | 3,420 | -31% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +14 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.