Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Score, evaluate, and iteratively improve any content or strategy using an auto-assembled panel of domain experts. Handles copy, sequences, landing pages, strategy docs, titles, charts, recruiting evaluations, or anything else that needs a quality gate. Recursively iterates until all scores hit 90+ (max 3 rounds). Use when asked to: "expert panel this", "score this", "rate these variants", "quality check this", "panel review", "which version is better", "expert score", "evaluate this copy/strateg
.claude/skills/evolution-foundation-mkt-quality-gate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 148% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 220% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 149% | 0% |
General-purpose scoring and iterative improvement engine. Auto-assembles the right experts for whatever is being evaluated, scores it, and loops until 90+.
Collect or infer from context:
If yes, note the source for feedback-to-source routing in Step 6.
If context is obvious from the conversation, don't ask — just proceed.
Build a panel of 7–10 experts tailored to the content type and domain.
experts/ directory for pre-built panels matchingthe content type. If an exact match exists (e.g., experts/linkedin.md for a LinkedIn post), use it as the base.
the specific industry or domain. Examples:
experts/humanizer.md. Weight: 1.5x. Non-negotiable.known rejection patterns from references/patterns.md (if present).
references/patterns.md exists, read it. If any patternsapply to this content type, brief the panel on them. Dock points for known-bad patterns.
List each expert with: Name, lens/focus, what they check.
Choose the appropriate rubric from scoring-rubrics/:
| Content type | Rubric file | |---|---| | Blog, social, email, newsletter, scripts | scoring-rubrics/content-quality.md | | Strategy, recommendations, analysis | scoring-rubrics/strategic-quality.md | | Landing pages, ads, CTAs | scoring-rubrics/conversion-quality.md | | Charts, data viz, infographics | scoring-rubrics/visual-quality.md | | Candidate evaluations | scoring-rubrics/evaluation-quality.md | | Other | Synthesize a rubric from the two closest matches |
Read the selected rubric file for detailed criteria and point allocation.
Target: 90/100 across all experts. Non-negotiable. Max 3 rounds.
## Round [N] — Score: [AVG]/100
| Expert | Score | Key Feedback |
|--------|-------|--------------|
| [Name] | [0-100] | [One-line rationale] |
| ... | ... | ... |
**Aggregate:** [weighted average — humanizer at 1.5x]
**Top 3 weaknesses:** [ranked]
**Changes made:** [specific edits addressing each weakness]Then the revised content/artifact.
holding it back.
When scoring multiple variants (A/B/C):
## 🏆 Result: [SCORE]/100 — [PASS ✅ | NEEDS WORK ⚠️]
[Final content/artifact here]
**Iterations:** [N] rounds
**Panel:** [Expert names, comma-separated]If variants: show winner first, then runner-up scores.
## 🏆 Winner: Variant [X] — [SCORE]/100
[Winning content]
### Runner-up scores
- Variant A: 87/100
- Variant B: 82/100
- Variant C: 91/100 ← WinnerShow full scoring rounds.
---
<details>
<summary>📊 Scoring History (N rounds)</summary>
[All round tables from Step 4]
</details>When the scored content came from another skill, generate a Source Improvement Brief:
## 🔁 Feedback for [Source Skill]
### What scored low
- [Pattern]: [Specific example from this content]
### Suggested skill improvements
- [Concrete change to the source skill's process/rubric/prompt]
### Patterns to add to source skill
- [Any recurring weakness that should become a rule]This brief can be used to update the source skill's SKILL.md or rubrics.
After the user approves or rejects panel output:
Note what worked. No action needed unless a new positive pattern emerges.
references/patterns.md using this format:markdown## [Pattern Name] - **Type:** rejection | preference | override - **Content types:** [which types this applies to] - **Rule:** [What to always/never do] - **Example:** [The specific instance that triggered this] - **Date:** [YYYY-MM-DD] - **Point dock:** [-N points when detected]
Every scoring round, check references/patterns.md against the content. Apply point docks before expert scoring begins. This means known-bad patterns are penalized even if individual experts miss them.
| File | Purpose | When to read | |---|---|---| | experts/humanizer.md | AI writing detection rubric (24 patterns) | Every scoring run | | experts/[domain].md | Pre-built expert panels for common domains | When domain matches | | scoring-rubrics/content-quality.md | Content scoring rubric | Content scoring | | scoring-rubrics/strategic-quality.md | Strategy scoring rubric | Strategy scoring | | scoring-rubrics/conversion-quality.md | Landing page/ad/CTA rubric | Conversion scoring | | scoring-rubrics/visual-quality.md | Chart/data viz/infographic rubric | Visual scoring | | scoring-rubrics/evaluation-quality.md | Candidate/assessment rubric | Eval scoring | | references/patterns.md | Learned rejection patterns | Every scoring run | | references/expert-assembly.md | Domain-expert examples for auto-assembly | When building unfamiliar panels |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 23,190 | 4,892 | -79% | 1 | 1 | 0% | 4,148 | 2,748 | -34% | 0 | 0 | — |
case-02 | fail→pass | 26,661 | 27,555 | +3% | 1 | 1 | 0% | 4,285 | 6,859 | +60% | 0 | 0 | — |
case-03 | fail→fail | 19,186 | 5,277 | -72% | 1 | 1 | 0% | 3,302 | 2,797 | -15% | 0 | 0 | — |
case-04 | fail→pass | 35,833 | 34,403 | -4% | 1 | 1 | 0% | 5,638 | 7,449 | +32% | 0 | 0 | — |
case-05 | fail→pass | 9,103 | 9,544 | +5% | 1 | 1 | 0% | 1,409 | 3,501 | +148% | 0 | 0 | — |
case-06 | fail→fail | 7,145 | 4,132 | -42% | 1 | 1 | 0% | 1,086 | 2,534 | +133% | 0 | 0 | — |
case-07 | fail→fail | 6,338 | 7,597 | +20% | 1 | 1 | 0% | 941 | 3,176 | +238% | 0 | 0 | — |
case-08 | fail→fail | 4,482 | 6,436 | +44% | 1 | 1 | 0% | 709 | 2,883 | +307% | 0 | 0 | — |
case-09 | fail→pass | 6,535 | 9,045 | +38% | 1 | 1 | 0% | 1,051 | 3,364 | +220% | 0 | 0 | — |
case-10 | fail→fail | 4,866 | 5,879 | +21% | 1 | 1 | 0% | 847 | 2,937 | +247% | 0 | 0 | — |
case-11 | fail→pass | 6,964 | 4,351 | -38% | 1 | 1 | 0% | 1,071 | 2,672 | +149% | 0 | 0 | — |
case-12 | fail→pass | 2,398 | 3,752 | +56% | 1 | 1 | 0% | 396 | 2,587 | +553% | 0 | 0 | — |
case-13 | fail→fail | 10,734 | 8,552 | -20% | 1 | 1 | 0% | 1,692 | 3,433 | +103% | 0 | 0 | — |
case-14 | fail→fail | 9,894 | 7,306 | -26% | 1 | 1 | 0% | 1,520 | 3,049 | +101% | 0 | 0 | — |
case-15 | fail→fail | 3,690 | 3,583 | -3% | 1 | 1 | 0% | 511 | 2,509 | +391% | 0 | 0 | — |
case-16 | fail→fail | 6,363 | 4,632 | -27% | 1 | 1 | 0% | 923 | 2,617 | +184% | 0 | 0 | — |
case-17 | fail→pass | 16,954 | 9,007 | -47% | 1 | 1 | 0% | 2,612 | 3,487 | +33% | 0 | 0 | — |
case-18 | fail→fail | 15,379 | 9,914 | -36% | 1 | 1 | 0% | 2,337 | 3,414 | +46% | 0 | 0 | — |
case-19 | fail→fail | 3,297 | 5,249 | +59% | 1 | 1 | 0% | 427 | 2,759 | +546% | 0 | 0 | — |
case-20 | pass→fail | 14,487 | 19,651 | +36% | 1 | 1 | 0% | 2,540 | 5,572 | +119% | 0 | 0 | — |
case-21 | pass→fail | 17,682 | 20,613 | +17% | 1 | 1 | 0% | 2,780 | 5,286 | +90% | 0 | 0 | — |
case-22 | pass→fail | 7,156 | 18,778 | +162% | 1 | 1 | 0% | 1,517 | 5,900 | +289% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases. 6 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.