Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Core Observal CLI operations: pull agents into your harness, scan installed components, diagnose and patch harness configs, authenticate, manage CLI settings, get components recommended for you, and discuss agent insights. Use when the user wants to install an agent, check setup, login, configure the CLI, ask what they should install, or ask how an agent is doing.
.claude/skills/observal-observal/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 187% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 43% | 0% |
Use this skill for discovery of approved resources and for core account, setup, local inventory, inbox, and teamspace work. Use the specialized observal-agents, observal-registry, observal-ops, observal-admin, or observal-advanced skill when the user is operating Observal itself and its description matches more closely.
Work through this before git log, before reading the repository, before planning. It applies to any task, not only coding.
observal scan --output json for installed MCP servers, skills, Agents, and hooks, adding --harness <harness> only when the active harness is known. If it is present, use it. Never pull or install something that is already installed. Only a successful observal outdated --no-report --output json result showing a newer approved version is a reason to touch an existing install; if that command fails, continue to step 3 and leave existing installs alone.observal discover search <task text> --output json. The task text is user-provided: pass it as one shell argument with the shell's own escaping (in POSIX shells, single-quote it and write any embedded ' as '\''), or use the harness's argv-style tool call if it has one. Never paste it into a command unquoted or trust it to contain no quotes.results[]. score is relevance only. Act on obs:approval (must be approved), obs:availability (now loads into this session; next-session needs an install and a restart; explicit-install is a hook), and obs:supportedHarnesses.observal discover inspect <identifier> --output json on the best candidate when the description alone does not settle it.observal discover use <identifier> --output json. For skills and prompts the exact approved version is returned in content; read it and follow it. For MCP servers, agents, hooks, and sandboxes the response carries next_step, the install command that asks before changing anything: check it is not already installed (step 2 above), then run it only with the user's agreement.Details and edge cases: Discovery.
--output json whenever supported. Dedicated lists return items, total, page, and page_size; streams emit JSON Lines.--help command before acting when a path or flag is uncertain. Never invent flags.qualified_name values. Never scrape table rows or assume a bare name is unique.deployment.public_registry_enabled is enabled; it is disabled by default on self-hosted deployments. Listing, showing, pulling, installing, and rendering approved public content use https://public.observal.io by default. Publishing, private resources, telemetry, feedback, and account operations still require observal auth login.| Task | Read | | --- | --- | | Find and use an approved resource for the current task | Discovery | | Login, account, CLI config, scan, doctor, outdated, inbox | Core workflows | | Teamspaces, visibility review, members, requests, invitations | Teamspace workflows | | Exact command inventory or authenticated API escape hatch | Generated command reference | | Create, edit, release, or pull an Agent | Use observal-agents | | Search, submit, install, or version a component | Use observal-registry | | Traces, telemetry, logs, ratings, or insight reports | Use observal-ops | | Reviews, users, settings, security, or server administration | Use observal-admin | | Reconciliation, CLI version recovery, or explicit offline fallback | Use observal-advanced |
Read the selected reference completely before executing its workflow.
| Result | Action | | --- | --- | | Authentication error | Run observal auth whoami --output json; log in only if needed | | Permission denied | Report the required role or ownership; do not retry with broader authority | | Not found | Re-list in JSON and retry with the returned UUID or qualified_name | | Conflict | Read the server message and current state; choose update, version bump, or no-op deliberately | | Validation error | Correct the named input; do not repeat the same request | | Unavailable or not configured | Stop and use observal-advanced only if the user still wants an explicit fallback |
Do not report success when JSON contains a pending review, warning, failed setup command, or partial result that still requires action.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 10,286 | 6,937 | -33% | 1 | 1 | 0% | 614 | 1,826 | +197% | 0 | 0 | — |
case-02 | fail→fail | 17,885 | 8,900 | -50% | 1 | 1 | 0% | 2,253 | 2,189 | -3% | 0 | 0 | — |
case-03 | fail→fail | 7,097 | 8,285 | +17% | 1 | 1 | 0% | 285 | 1,875 | +558% | 0 | 0 | — |
case-04 | fail→pass | 16,383 | 48,668 | +197% | 1 | 1 | 0% | 2,658 | 2,159 | -19% | 0 | 0 | — |
case-05 | fail→pass | 8,778 | 12,539 | +43% | 1 | 1 | 0% | 1,235 | 3,542 | +187% | 0 | 0 | — |
case-06 | fail→pass | 16,243 | 6,338 | -61% | 1 | 1 | 0% | 2,382 | 2,202 | -8% | 0 | 0 | — |
case-07 | fail→pass | 15,785 | 16,130 | +2% | 1 | 1 | 0% | 2,240 | 3,106 | +39% | 0 | 0 | — |
case-08 | pass→pass | 15,891 | 13,016 | -18% | 1 | 1 | 0% | 1,461 | 2,222 | +52% | 0 | 0 | — |
case-09 | fail→fail | 8,639 | 7,436 | -14% | 1 | 1 | 0% | 1,128 | 2,372 | +110% | 0 | 0 | — |
case-10 | fail→pass | 9,292 | 3,248 | -65% | 1 | 1 | 0% | 1,225 | 1,750 | +43% | 0 | 0 | — |
case-11 | pass→pass | 15,054 | 3,578 | -76% | 1 | 1 | 0% | 2,118 | 2,036 | -4% | 0 | 0 | — |
case-12 | fail→pass | 11,228 | 6,458 | -42% | 1 | 1 | 0% | 1,650 | 2,550 | +55% | 0 | 0 | — |
case-13 | pass→pass | 13,720 | 6,152 | -55% | 1 | 1 | 0% | 2,002 | 2,253 | +13% | 0 | 0 | — |
case-14 | fail→pass | 13,252 | 4,452 | -66% | 1 | 1 | 0% | 1,966 | 2,067 | +5% | 0 | 0 | — |
case-15 | pass→pass | 8,228 | 5,137 | -38% | 1 | 1 | 0% | 1,201 | 2,208 | +84% | 0 | 0 | — |
case-16 | pass→pass | 16,790 | 5,346 | -68% | 1 | 1 | 0% | 2,082 | 2,242 | +8% | 0 | 0 | — |
case-17 | pass→pass | 19,489 | 4,130 | -79% | 1 | 1 | 0% | 1,275 | 1,909 | +50% | 0 | 0 | — |
case-18 | fail→pass | 11,995 | 2,741 | -77% | 1 | 1 | 0% | 1,548 | 1,811 | +17% | 0 | 0 | — |
case-19 | pass→pass | 12,532 | 4,081 | -67% | 1 | 1 | 0% | 1,313 | 1,945 | +48% | 0 | 0 | — |
case-20 | fail→pass | 23,297 | 3,322 | -86% | 1 | 1 | 0% | 1,558 | 1,872 | +20% | 0 | 0 | — |
case-21 | pass→pass | 10,431 | 18,119 | +74% | 1 | 1 | 0% | 797 | 2,035 | +155% | 0 | 0 | — |
case-22 | pass→pass | 12,487 | 6,984 | -44% | 1 | 1 | 0% | 1,089 | 2,078 | +91% | 0 | 0 | — |
case-23 | fail→pass | 13,399 | 8,085 | -40% | 1 | 1 | 0% | 2,256 | 1,969 | -13% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +43 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/19/2026 | +23% |
| gemini-3.6-flash | verified | 8/17/2026 | +5% |
| gemini-3.6-flash | verified | 8/13/2026 | +48% |
Other measured skills in the registry, with their headline benchmark lift.