Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Auto-detect work context (Dev vs Knowledge) — use to tailor workflows based on current task type
.claude/skills/hashgraph-online-skill-context-detection/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 217% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 287% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 143% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 217% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 30% | 0% |
> Host: Codex CLI — This skill was designed for Claude Code and adapted for Codex. > Cross-reference commands use installed skill names in Codex rather than /octo:* slash commands. > Use the active Codex shell and subagent tools. Do not claim a provider, model, or host subagent is available until the current session exposes it. > For host tool equivalents, see skills/blocks/codex-host-adapter.md.
This skill provides automatic context detection to determine whether the user is working in a Development context (code-focused) or Knowledge context (research/strategy-focused). This replaces the manual /octo:km toggle with intelligent auto-detection.
When a workflow skill activates, detect context using these signals:
If user has explicitly set mode via /octo:km on or /octo:km off, respect that setting.
bash# Check if knowledge mode is explicitly set if [[ -f ~/.claude-octopus/config/knowledge-mode ]]; then EXPLICIT_MODE=$(cat ~/.claude-octopus/config/knowledge-mode) if [[ "$EXPLICIT_MODE" == "on" ]]; then echo "knowledge" exit 0 elif [[ "$EXPLICIT_MODE" == "off" ]]; then echo "dev" exit 0 fi fi # If "auto" or not set, proceed with auto-detection
Knowledge Context Indicators (check prompt for these terms):
Dev Context Indicators (check prompt for these terms):
Scoring:
Dev Project Indicators:
package.json, Cargo.toml, go.mod, pyproject.toml, pom.xmlsrc/, lib/, app/ directories with code files.ts, .js, .py, .go, .rs, .javaKnowledge Project Indicators:
docs/, research/, strategy/, reports/ directories.md, .docx, .pdf, .pptxIf signals are ambiguous or equal:
Return detected context as a structured object for use by workflow skills:
json{ "context": "dev" | "knowledge", "confidence": "high" | "medium" | "low", "signals": { "prompt_indicators": ["API", "endpoint", "database"], "project_type": "node_typescript", "explicit_override": false } }
| Aspect | Dev Context | Knowledge Context | |--------|-------------|-------------------| | Research Focus | Technical implementation, library comparison, code patterns | Market analysis, academic synthesis, competitive research | | Primary Agents | Codex (implementation), Antigravity (ecosystem) | Antigravity (analysis), research-synthesizer | | Output Format | Code examples, API comparisons, tech recommendations | Reports, frameworks, strategic recommendations | | Visual Banner | 🔍 [Dev] Discover Phase: Technical research | 🔍 [Knowledge] Discover Phase: Strategic research |
| Aspect | Dev Context | Knowledge Context | |--------|-------------|-------------------| | Build Focus | Code generation, implementation, architecture | PRDs, strategy docs, presentations | | Primary Agents | Codex (code), backend-architect, tdd-orchestrator | product-writer, strategy-analyst, exec-communicator | | Output Format | Source files, tests, migrations | Documents, frameworks, action plans | | Visual Banner | 🛠️ [Dev] Develop Phase: Building code | 🛠️ [Knowledge] Develop Phase: Building deliverables |
| Aspect | Dev Context | Knowledge Context | |--------|-------------|-------------------| | Review Focus | Code quality, security, performance | Document quality, argument strength, completeness | | Primary Agents | code-reviewer, security-auditor | exec-communicator, strategy-analyst | | Quality Gates | OWASP, test coverage, maintainability | Evidence quality, clarity, actionability | | Visual Banner | ✅ [Dev] Deliver Phase: Code review | ✅ [Knowledge] Deliver Phase: Document review |
When context is detected, update the visual banner to show context:
Dev Context:
🐙 **CLAUDE OCTOPUS ACTIVATED** - Multi-provider research mode
🔍 [Dev] Discover Phase: Researching OAuth implementation patterns
Providers:
🔴 Codex CLI - Technical implementation analysis
🟡 Antigravity CLI - Ecosystem and library comparison
🔵 Claude - Strategic synthesisKnowledge Context:
🐙 **CLAUDE OCTOPUS ACTIVATED** - Multi-provider research mode
🔍 [Knowledge] Discover Phase: Researching market entry strategies
Providers:
🔴 Codex CLI - Data analysis and modeling
🟡 Antigravity CLI - Market and competitive research
🔵 Claude - Strategic synthesisEach flow skill should:
markdownWhen this skill activates: 1. **Detect context** - Analyze user's prompt for knowledge vs dev indicators - Check project type (code repo vs doc-heavy) - Check for explicit override (~/.claude-octopus/config/knowledge-mode) - Determine: "dev" or "knowledge" with confidence level 2. **Show context-aware banner** ``` 🐙 **CLAUDE OCTOPUS ACTIVATED** - Multi-provider [research|implementation|validation] mode [Phase Emoji] [Context] [Phase Name]: [Description] Detected Context: [Dev|Knowledge] (confidence: [high|medium|low]) ``` 3. **Execute workflow with context-appropriate behavior** - Frame prompts for external providers based on context - Select appropriate synthesis approach - Apply context-specific quality gates
Users can still explicitly set context when auto-detection is wrong:
bash# Force knowledge mode /octo:km on # Force dev mode /octo:km off # Return to auto-detection /octo:km auto
When explicit override is set, context detection respects it until user resets to "auto".
When confidence is "low", consider briefly mentioning the detected context to user: > "I detected this as a dev/knowledge] task. If that's wrong, you can use /octo:km to override."
To verify context detection is working:
/octo:km on set, ask "octo research API patterns" → Should use Knowledge Context (explicit override)When detecting the user's work stage, surface relevant command suggestions:
| Detected Context | Suggestion | |-----------------|------------| | Brainstorming / exploring ideas | Consider /octo:brainstorm for structured ideation | | Reviewing a plan or strategy | Consider /octo:plan for strategic planning | | Debugging errors or failures | Consider /octo:debug for systematic investigation | | Writing or running tests | Consider /octo:tdd for test-driven development | | Code review before merge | Use Claude-native /review for ordinary review; suggest /octo:review for multi-AI escalation | | Ready to deploy or ship | Consider /octo:deliver for quality-gated delivery | | Researching a topic | Consider /octo:research for multi-source synthesis | | Working on security | Use Claude-native /security-review for ordinary security review; suggest /octo:security for escalated OWASP or adversarial audit |
Suggestions should be non-intrusive, appended as a brief note:
💡 Tip: You appear to be debugging — `/octo:debug` provides systematic investigation with multi-AI support.OCTO_PROACTIVE_SUGGESTIONS=off in .claude-octopus/preferences.jsonOCTO_PROACTIVE_SUGGESTIONS=onUsers who previously opted out can re-enable suggestions at any time:
~/.claude-octopus/preferences.json and set OCTO_PROACTIVE_SUGGESTIONS to onDetect work stage from:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,050 | 16,401 | +225% | 1 | 1 | 0% | 990 | 2,831 | +186% | 0 | 0 | — |
case-02 | fail→fail | 4,043 | 16,460 | +307% | 1 | 1 | 0% | 720 | 2,979 | +314% | 0 | 0 | — |
case-03 | fail→fail | 10,922 | 11,621 | +6% | 1 | 1 | 0% | 1,135 | 3,047 | +168% | 0 | 0 | — |
case-04 | fail→fail | 4,532 | 15,805 | +249% | 1 | 1 | 0% | 254 | 2,847 | +1021% | 0 | 0 | — |
case-05 | pass→pass | 58,782 | 35,436 | -40% | 1 | 1 | 0% | 8,217 | 7,807 | -5% | 0 | 0 | — |
case-06 | pass→fail | 4,835 | 10,533 | +118% | 1 | 1 | 0% | 870 | 2,771 | +219% | 0 | 0 | — |
case-07 | fail→fail | 8,509 | 12,697 | +49% | 1 | 1 | 0% | 1,412 | 2,938 | +108% | 0 | 0 | — |
case-08 | fail→pass | 8,957 | 3,373 | -62% | 1 | 1 | 0% | 1,010 | 3,204 | +217% | 0 | 0 | — |
case-09 | fail→pass | 9,118 | 3,083 | -66% | 1 | 1 | 0% | 789 | 3,053 | +287% | 0 | 0 | — |
case-10 | fail→pass | 11,563 | 7,806 | -32% | 1 | 1 | 0% | 1,273 | 3,090 | +143% | 0 | 0 | — |
case-11 | fail→pass | 5,570 | 9,353 | +68% | 1 | 1 | 0% | 1,061 | 3,365 | +217% | 0 | 0 | — |
case-17 | fail→pass | 27,429 | 5,020 | -82% | 1 | 1 | 0% | 2,277 | 2,962 | +30% | 0 | 0 | — |
case-12 | fail→pass | 12,086 | 5,284 | -56% | 1 | 1 | 0% | 1,027 | 3,527 | +243% | 0 | 0 | — |
case-13 | fail→pass | 9,100 | 4,754 | -48% | 1 | 1 | 0% | 1,561 | 2,956 | +89% | 0 | 0 | — |
case-14 | fail→pass | 11,474 | 2,737 | -76% | 1 | 1 | 0% | 1,196 | 2,948 | +146% | 0 | 0 | — |
case-15 | fail→pass | 6,531 | 2,911 | -55% | 1 | 1 | 0% | 1,135 | 2,821 | +149% | 0 | 0 | — |
case-16 | fail→pass | 13,745 | 7,155 | -48% | 1 | 1 | 0% | 2,484 | 2,819 | +13% | 0 | 0 | — |
case-18 | fail→pass | 11,121 | 2,538 | -77% | 1 | 1 | 0% | 1,867 | 2,979 | +60% | 0 | 0 | — |
case-19 | pass→pass | 7,289 | 2,381 | -67% | 1 | 1 | 0% | 371 | 2,957 | +697% | 0 | 0 | — |
case-20 | fail→pass | 4,492 | 2,768 | -38% | 1 | 1 | 0% | 806 | 2,973 | +269% | 0 | 0 | — |
case-21 | fail→pass | 14,005 | 3,744 | -73% | 1 | 1 | 0% | 2,588 | 3,136 | +21% | 0 | 0 | — |
case-22 | fail→pass | 28,524 | 6,021 | -79% | 1 | 1 | 0% | 1,091 | 3,662 | +236% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 15 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.