Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Core Observal CLI operations: pull agents into your IDE, scan installed components, diagnose and patch IDE configs, authenticate, and manage CLI settings. Use when the user wants to install an agent, check their IDE setup, login, or configure the CLI.
.claude/skills/bilal140202-observal/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 161% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 79% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 165% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 25% | 0% |
--prompt and --description values to avoid shell quoting issues.observal auth status first. Other commands surface auth problems clearly on their own.<command> --help first. Never guess flag names.--output json on every list/show command. It is stable and machine readable.--yes / -y on destructive commands so they do not block on a confirmation prompt.--update for in-place edits, --bump for versioned releases.Connection failed or Not configured.OTEL_* or CLAUDE_CODE_ENABLE_TELEMETRY environment variables. Telemetry flows through observal-shim and session push hooks only.Install an agent's full config (rules, MCP servers, hooks, skills) into a local IDE.
bashobserval agent pull AGENT_NAME --ide kiro --no-prompt --dir .
Flags:
--ide (required): claude-code, kiro, cursor, gemini-cli, vscode, codex, copilot, copilot-cli, opencode--version <semver>: install a specific version (e.g. 1.2.0). Omit for latest.--scope user|project: install scope (Claude Code, Kiro, Gemini only)--model <name> or --model <ide>=<name>: override saved model (repeatable)--tools t1,t2: Claude Code tool whitelist--dry-run: preview file writes without touching disk--no-prompt: skip interactive confirmation--dir <path>: target directory (default: current)Merge behavior: MCP configs are merged with existing IDE config files, not overwritten. Existing user entries are preserved.
Version pinning: When --version is specified, the exact content from that version is installed. The lockfile (~/.observal/lockfile.json) records the pin. If another agent depends on the same component at a different version, a warning is displayed.
If the user did not specify an IDE, ask which one before running.
Check for newer versions of installed agents and components.
bashobserval outdated observal outdated --ide claude-code observal outdated --output json
Reads ~/.observal/lockfile.json and compares each pinned version against the registry's latest. Reports a table of outdated items with current vs latest version.
Read-only inventory of installed components across all detected IDEs. Never modifies any file.
bashobserval scan observal scan --ide kiro observal scan --ide claude-code
Reports: detected IDEs, MCP servers (with shimmed status), skills, hooks, agents, and unregistered components.
Diagnose only. Does not fix anything.
bashobserval doctor
Reports: Observal config validity, server reachability, hook installation status per IDE, skill presence. Exits non-zero if issues found.
Apply instrumentation. Run with --dry-run first when the user is unsure.
bashobserval doctor patch --all --all-ides --dry-run observal doctor patch --all --all-ides observal doctor patch --hook --shim --ide kiro observal doctor patch --all --ide claude-code observal doctor patch --hook --all-ides observal doctor patch --shim --all-ides
Required: at least one of --hook / --shim / --all, AND at least one of --all-ides / --ide. Creates timestamped backups before modifying any file.
Remove Observal-managed hooks and env vars from IDE configs. Leaves user content untouched.
bashobserval doctor cleanup --dry-run observal doctor cleanup observal doctor cleanup --ide kiro
bashobserval auth login observal auth login --server https://observal.example.com observal auth login --sso observal auth login --email me@x.com --password '...' observal auth whoami --output json observal auth status observal auth logout observal auth change-password observal auth set-username new-handle
On a fresh server, auth login auto-bootstraps an admin from localhost (no prompts needed).
bashobserval config show observal config path observal config set output json observal config set server_url https://observal.example.com observal config aliases observal config alias MY_AGENT abc-123
| Error | Action | |-------|--------| | Connection failed | Server unreachable. Use the observal-advanced skill's Local Fallback procedure | | Not configured / No server | Run observal auth login | | 403 Forbidden | Check observal auth whoami; user lacks required role | | 404 Not found | Verify name with observal agent list --output json |
For every CLI invocation, format your response:
For full command reference, read references/commands.md. For agent creation use the observal-agents skill. For registry operations use observal-registry. For observability use observal-ops. For admin tasks use observal-admin.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | pass→fail | 5,176 | 6,052 | +17% | 1 | 1 | 0% | 990 | 1,846 | +86% | 0 | 0 | — |
case-01 | fail→pass | 3,408 | 1,866 | -45% | 1 | 1 | 0% | 686 | 1,791 | +161% | 0 | 0 | — |
case-02 | fail→pass | 12,342 | 5,466 | -56% | 1 | 1 | 0% | 1,754 | 1,685 | -4% | 0 | 0 | — |
case-03 | fail→pass | 5,062 | 4,026 | -20% | 1 | 1 | 0% | 1,030 | 1,846 | +79% | 0 | 0 | — |
case-04 | fail→fail | 6,330 | 5,264 | -17% | 1 | 1 | 0% | 1,158 | 1,801 | +56% | 0 | 0 | — |
case-05 | fail→pass | 4,114 | 2,854 | -31% | 1 | 1 | 0% | 744 | 1,972 | +165% | 0 | 0 | — |
case-07 | fail→fail | 9,491 | 4,622 | -51% | 1 | 1 | 0% | 1,714 | 1,734 | +1% | 0 | 0 | — |
case-08 | fail→pass | 7,656 | 1,668 | -78% | 1 | 1 | 0% | 1,388 | 1,729 | +25% | 0 | 0 | — |
case-09 | fail→pass | 7,048 | 2,749 | -61% | 1 | 1 | 0% | 1,132 | 1,921 | +70% | 0 | 0 | — |
case-10 | fail→pass | 11,021 | 2,067 | -81% | 1 | 1 | 0% | 2,114 | 1,794 | -15% | 0 | 0 | — |
case-11 | fail→pass | 12,818 | 3,912 | -69% | 1 | 1 | 0% | 2,595 | 1,674 | -35% | 0 | 0 | — |
case-12 | fail→pass | 6,455 | 8,599 | +33% | 1 | 1 | 0% | 1,055 | 2,277 | +116% | 0 | 0 | — |
case-13 | fail→pass | 3,455 | 2,032 | -41% | 1 | 1 | 0% | 610 | 1,749 | +187% | 0 | 0 | — |
case-14 | pass→pass | 7,025 | 4,003 | -43% | 1 | 1 | 0% | 1,323 | 1,776 | +34% | 0 | 0 | — |
case-15 | fail→fail | 6,470 | 3,059 | -53% | 1 | 1 | 0% | 1,397 | 1,892 | +35% | 0 | 0 | — |
case-16 | fail→pass | 9,386 | 5,672 | -40% | 1 | 1 | 0% | 1,347 | 2,375 | +76% | 0 | 0 | — |
case-17 | fail→pass | 8,278 | 1,716 | -79% | 1 | 1 | 0% | 1,309 | 1,607 | +23% | 0 | 0 | — |
case-18 | pass→fail | 8,345 | 7,794 | -7% | 1 | 1 | 0% | 1,512 | 1,889 | +25% | 0 | 0 | — |
case-19 | fail→pass | 12,460 | 3,062 | -75% | 1 | 1 | 0% | 2,039 | 1,704 | -16% | 0 | 0 | — |
case-20 | fail→fail | 9,750 | 5,692 | -42% | 1 | 1 | 0% | 1,469 | 1,885 | +28% | 0 | 0 | — |
case-21 | fail→pass | 5,915 | 2,859 | -52% | 1 | 1 | 0% | 956 | 1,807 | +89% | 0 | 0 | — |
case-22 | fail→fail | 11,656 | 1,859 | -84% | 1 | 1 | 0% | 1,767 | 1,706 | -3% | 0 | 0 | — |
case-23 | fail→fail | 5,370 | 5,079 | -5% | 1 | 1 | 0% | 824 | 1,708 | +107% | 0 | 0 | — |
case-24 | fail→fail | 6,557 | 3,069 | -53% | 1 | 1 | 0% | 1,268 | 1,957 | +54% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 19 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 19 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.