Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing or auditing ACM CCS experiments, attack demonstrations, adaptive-attack defense evaluations, security measurements, baselines, overhead and cost reporting, ablations, and claim-to-evidence fit, with emphasis on evidence that survives an adversarial program committee rather than leaderboard wins.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-01 | ✓→✓ | = Same ✓ | -30% | 0% |
Use this before submission when the attack demonstration, defense evaluation, or measurement story is not yet locked.
coverage number, a false-positive/false-negative table, or a measurement dataset.
platform, configuration) and report the resource cost to the attacker.
and report performance overhead, memory cost, and any compatibility breakage.
and blind spots, and ground-truth checks against known cases.
mismatch between the threat model and the tested configuration.
benchmark. One clean end-to-end exploit against a real target outweighs a table of micro-benchmarks.
defense, and a deployment-cost measurement. Missing the middle element is the classic CCS defense reject.
model. A defense claimed for production but tested only on a toy in a lab invites the relevance question.
| Security claim | Matching evidence | Reject pattern avoided | |---|---|---| | Exploit is practical | End-to-end run on named target with attacker cost | "Works only in a lab against a strawman" | | Defense stops the attack | Detection/prevention rate on the original attack | "No numbers, only a design argument" | | Defense resists adaptation | Adaptive attacker with defense knowledge, degraded results | "Only the non-adaptive attack was tried" | | Deployment is feasible | Overhead, memory, compatibility on a realistic workload | "Security claimed, cost never measured" |
Suppose the paper proposes a fine-grained CFI scheme. The matching plan: reproduce a known code-reuse attack and show it blocked; construct an adaptive attacker that respects the CFI policy and search for surviving gadget chains; then measure runtime overhead and binary-size growth on a standard benchmark suite. Every claim ties to a numbered table, and the adaptive result is reported even when it dents the headline.
overhead rather than vague "negligible cost" language.
text[Experiment readiness] strong / adequate / weak [Claim -> evidence map] <claim: exploit run / overhead table / measurement> [Missing security evidence] <adaptive attack / baseline / cost / validation> [Threat-model mismatch] <where the setup breaks the stated model> [Decision-critical next run] <one experiment>
Other measured skills in the registry, with their headline benchmark lift.