Install any skill in seconds. Free to start, no credit card required.
Get Started Free →[omh] Hermes Agent Evaluation workflow: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor benchmark, compare agents, compare codex claude.
.claude/skills/rlaope-omh-agent-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 157% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -5% | 0% |
This is a Hermes-native agent-evaluation workflow skill.
agent-evaluation gives OMH a way to improve executor choice empirically, not by vibes, while preserving executor-neutral product language across Codex, Claude Code, Hermes, and generic runtimes.
executor-runtime-readiness.workflow-learning.ultraperf.Good example:
Bad example:
achievements, workspace-audit, production-audit, automation-blueprint, github-event-ops, buzz, agent-board, gateway-intent-card, +34 more) - schedules, status, health, and ops review.oh-my-hermes or name the adjacent workflow.omh-routing/references/skill-common-rail.md.Use when Hermes should design or summarize a fair comparison of Codex, Claude Code, Hermes coding, or generic executors for a bounded task set.
Strong routing signals: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor benchmark, compare agents, compare codex claude, agent tournament, which agent is better, 에이전트 평가, 에이전트 비교, 실행자 평가, 코덱스 클로드 비교
Category: operations Phase: agent-evaluation Hermes role: operator Quality tier: agent-eval-gated Reasoning demand: light
Quality bar:
Handoff policy:
Keep evaluation design and scoring in Hermes. Actual executor runs, costs, timings, tool calls, code edits, and review results must come from observed runtime or supplied artifacts.
Required inputs:
Expected outputs:
Artifact expectations:
Safety rules:
Preferred harness for this skill: agent-evaluation.
shomh runtime record --skill agent-evaluation --harness agent-evaluation --status started
Record observed delegation results; otherwise return not_available or not_observed. Prepared OMH routing is not execution, review, CI, merge-readiness, or merge evidence.
Preserve workflow intent and stop conditions; verify before claiming completion.
Use Hermes-native subagent/delegation features when available: native subagents -> Hermes delegation when available, otherwise sequential lanes.
Shared product, compatibility, topology, memory, harness, and execution rules: omh-routing/references/skill-common-rail.md. Load it when applicable; otherwise name an unavailable capability.
Other measured skills in the registry, with their headline benchmark lift.