Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audit and remediate Entroly's MCP marketplace quality with evidence, adversarial validation, and no score gaming.
.claude/skills/juyterman1000-entroly-lobehub-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 118% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 2% | 0% |
Audit Entroly's complete LobeHub MCP score at:
https://lobehub.com/mcp/juyterman1000-entroly?activeTab=score
Act as a senior open-source product architect, MCP engineer, security engineer, Rust/Python/TypeScript developer, and release engineer.
index state, or genuinely missing capability.
evidence, fake benchmarks, or unnecessary features merely to game a score.
verification control plane for AI agents.
security controls, documentation, and reproducible evidence.
validation paths.
verify the public LobeHub page after its index refresh.
The current public LobeHub implementation assigns 100 total points:
| Criterion | Weight | Required | | --- | ---: | :---: | | Claimed listing | 4 | No | | Non-manual deployment | 12 | No | | Any deployment | 15 | Yes | | Detected license | 8 | No | | MCP prompts | 8 | No | | README | 10 | Yes | | MCP resources | 8 | No | | MCP tools | 15 | Yes | | Runtime validation | 20 | Yes |
All required criteria must pass. With all required criteria present, 80% or higher is grade A, 60-79% is grade B, and lower is grade F.
Maintain a table for each run:
| Criterion | Public observation | Repository evidence | Classification | Action | Test | External verification | | --- | --- | --- | --- | --- | --- | --- |
Never infer an external success from local code alone. Mark external-only results as pending, blocked, or confirmed with a direct artifact, registry response, public page, or screenshot.
registries, official specifications, and executable protocol probes.
an unproved assumption; do not disguise a blocked route as progress.
understand, validate, or safely use Entroly.
secret leakage, prompt injection, protocol compatibility, clean installs, and stale metadata.
CI matrix.
Reject any candidate remediation that:
The task is complete only when:
blocker;
external indexing.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 12,715 | 8,499 | -33% | 1 | 1 | 0% | 2,284 | 2,675 | +17% | 0 | 0 | — |
case-05 | pass→pass | 19,670 | 14,338 | -27% | 1 | 1 | 0% | 3,718 | 4,032 | +8% | 0 | 0 | — |
case-01 | fail→fail | 5,280 | 6,691 | +27% | 1 | 1 | 0% | 286 | 1,691 | +491% | 0 | 0 | — |
case-02 | fail→fail | 28,119 | 20,963 | -25% | 1 | 1 | 0% | 3,026 | 1,413 | -53% | 0 | 0 | — |
case-03 | fail→fail | 24,204 | 4,898 | -80% | 1 | 1 | 0% | 4,402 | 1,349 | -69% | 0 | 0 | — |
case-06 | fail→pass | 8,963 | 7,760 | -13% | 1 | 1 | 0% | 1,744 | 2,688 | +54% | 0 | 0 | — |
case-07 | fail→pass | 8,521 | 2,869 | -66% | 1 | 1 | 0% | 1,705 | 1,756 | +3% | 0 | 0 | — |
case-08 | fail→pass | 8,855 | 5,993 | -32% | 1 | 1 | 0% | 1,545 | 2,096 | +36% | 0 | 0 | — |
case-09 | fail→pass | 22,245 | 5,761 | -74% | 1 | 1 | 0% | 939 | 2,049 | +118% | 0 | 0 | — |
case-10 | fail→pass | 8,173 | 1,748 | -79% | 1 | 1 | 0% | 1,359 | 1,385 | +2% | 0 | 0 | — |
case-11 | fail→pass | 10,467 | 5,159 | -51% | 1 | 1 | 0% | 1,668 | 2,037 | +22% | 0 | 0 | — |
case-12 | pass→pass | 9,027 | 7,884 | -13% | 1 | 1 | 0% | 1,439 | 2,395 | +66% | 0 | 0 | — |
case-13 | fail→pass | 7,539 | 5,794 | -23% | 1 | 1 | 0% | 1,176 | 2,135 | +82% | 0 | 0 | — |
case-14 | pass→pass | 11,946 | 6,214 | -48% | 1 | 1 | 0% | 1,845 | 2,061 | +12% | 0 | 0 | — |
case-19 | pass→pass | 12,819 | 4,372 | -66% | 1 | 1 | 0% | 2,143 | 1,847 | -14% | 0 | 0 | — |
case-15 | pass→pass | 12,922 | 14,243 | +10% | 1 | 1 | 0% | 1,999 | 3,347 | +67% | 0 | 0 | — |
case-16 | fail→pass | 11,387 | 2,350 | -79% | 1 | 1 | 0% | 1,951 | 1,457 | -25% | 0 | 0 | — |
case-17 | fail→pass | 12,454 | 3,937 | -68% | 1 | 1 | 0% | 2,017 | 1,778 | -12% | 0 | 0 | — |
case-18 | fail→pass | 7,120 | 2,785 | -61% | 1 | 1 | 0% | 1,086 | 1,473 | +36% | 0 | 0 | — |
case-20 | pass→pass | 13,193 | 7,501 | -43% | 1 | 1 | 0% | 1,931 | 2,303 | +19% | 0 | 0 | — |
case-21 | fail→pass | 6,339 | 1,451 | -77% | 1 | 1 | 0% | 1,042 | 1,320 | +27% | 0 | 0 | — |
case-22 | fail→pass | 11,613 | 5,715 | -51% | 1 | 1 | 0% | 1,918 | 1,996 | +4% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.