Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Screening Assistant - AI-PRISMA 6-dimension screening with Groq LLM (100x cheaper) Supports two project types with different confidence thresholds Use when: screening papers, PRISMA screening, inclusion/exclusion criteria Triggers: screen papers, PRISMA screening, inclusion criteria, exclusion criteria, AI screening
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 905% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 65% | 0% |
diverga_check_prerequisites("i2") → must return approved: true If not approved → AskUserQuestion for each missing checkpoint (see .claude/references/checkpoint-templates.md)
diverga_mark_checkpoint("SCH_SCREENING_CRITERIA", decision, rationale)Read .research/decision-log.yaml directly to verify prerequisites. Conversation history is last resort.
Agent ID: I2 Category: I - Systematic Review Automation Tier: MEDIUM (Sonnet) Icon: 📋✅
Executes AI-assisted PRISMA 2020 screening using a 6-dimension rubric. Leverages Groq LLM for 100x cost reduction compared to Claude, while maintaining screening quality. Supports two project types with different confidence thresholds.
| Provider | Model | Cost per 100 papers | Quality | |----------|-------|---------------------|---------| | Groq (Default) | llama-3.3-70b | $0.01 | Excellent | | Groq | qwen-qwq-32b | $0.008 | Good | | Claude | claude-haiku-4-5 | $0.15 | Excellent | | Claude | claude-sonnet-3-5 | $0.45 | Best | | Ollama | llama3.2:70b | $0 | Good (local) |
Recommendation: Use Groq for screening. Switch to Claude only for complex edge cases.
yamlRequired: - project_path: "string" - research_question: "string" - project_type: "enum[knowledge_repository, systematic_review]" Optional: - llm_provider: "enum[groq, claude, ollama]" - custom_criteria: "object" - max_workers: "int" - batch_size: "int"
yamlmain_output: stage: "prisma_screening" project_type: "string" threshold: "int" llm_provider: "string" model: "string" results: total_screened: "int" auto_included: "int" auto_excluded: "int" human_review: "int" cost: input_tokens: "int" output_tokens: "int" total_cost: "string" output_files: relevant_papers: "string" excluded_papers: "string" human_review: "string"
Before executing screening, I2 MUST:
AI-PRISMA 6-Dimension Screening Criteria
Project Type: {knowledge_repository | systematic_review} Threshold: {50% | 90%} confidence
Scoring Rubric:
Total Score Range: -20 to 50 points
Decision Rules:
Do you approve these criteria?
bash# Project path (set to your working directory) cd "$(pwd)" # Set LLM provider (v1.2.6: Groq default) export LLM_PROVIDER=groq export GROQ_API_KEY={api_key} # Execute screening python scripts/03_screen_papers.py \ --project {project_path} \ --question "{research_question}" \ --max-workers 8 \ --batch-size 50
I2 validates AI evidence quotes against abstracts:
pythondef validate_evidence_grounding(quotes, abstract): """Flag potential hallucinations""" for quote in quotes: if quote.lower() not in abstract.lower(): return False, "FLAGGED: Potential hallucination" return True, None
Papers with hallucinated evidence are routed to human review.
| Keywords (EN) | Keywords (KR) | Action | |---------------|---------------|--------| | screen papers, PRISMA screening | 논문 스크리닝, 선별 | Activate I2 | | inclusion criteria, exclusion | 포함 기준, 제외 기준 | Activate I2 | | AI screening, automated screening | AI 스크리닝 | Activate I2 |
I2 can call B2-evidence-quality-appraiser for deeper quality assessment:
pythonTask( subagent_type="diverga:b2", model="sonnet", prompt=""" Assess quality of included papers using: - Risk of Bias (RoB) for RCTs - Newcastle-Ottawa for observational - GRADE for overall evidence quality """ )
yamlrequires: ["I1-paper-retrieval-agent"] sequential_next: ["I3-rag-builder"] parallel_compatible: ["B2-evidence-quality-appraiser"]
Other measured skills in the registry, with their headline benchmark lift.