Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Manage GPU compute jobs on the Qizhi (启智) platform using qzcli — a kubectl-style CLI tool. Use when user says "qzcli", "启智平台", "submit job", "stop job", "查计算组", "avail", "list jobs", "batch submit", or needs to manage distributed training jobs on a Qizhi instance.
.claude/skills/wanshuiyin-qzcli/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 115% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 63% | 0% |
A kubectl/docker-style CLI for managing GPU compute jobs on the Qizhi (启智) platform.
GitHub: tianyilt/qzcli_tool
Qizhi is the scheduler-cluster shape of ../shared-references/compute-env-contract.md: images are built OFF-platform and referenced at submit time, so the declarative env spec + env:<name>@<specHash> ledger (.aris/compute/qizhi.md) is what keeps "which image has which stack" answerable. Run the kernel witness inside a submitted job (not on the login side) before trusting an image for a long run.
bashpip install rich requests prompt_toolkit mcp git clone https://github.com/tianyilt/qzcli_tool cd qzcli_tool && pip install -e .
To use qzcli as an MCP tool directly from Claude Code or Codex:
bash# Claude Code claude mcp add qzcli -- qzcli-mcp # Codex codex mcp add qzcli -- qzcli-mcp
Credentials are read in this priority order: CLI args > --password-stdin > env vars > QZCLI_ENV_FILE (.env) > ~/.qzcli/config.json > interactive input
bash# Option A: env file (recommended) mkdir -p ~/.qzcli cat > ~/.qzcli/.env <<'EOF' QZCLI_USERNAME="your_username" QZCLI_PASSWORD="your_password" EOF # Option B: environment variables export QZCLI_USERNAME="your_username" export QZCLI_PASSWORD="your_password" export QZCLI_API_URL="https://qz.yourorg.edu.cn"
Config files are stored in ~/.qzcli/: config.json, .cookie, resources.json, jobs.json.
bash# 1. Login qzcli login # 2. Discover and cache workspaces/compute groups (run once, re-run after joining new workspaces) qzcli res -u # 3. Check available nodes qzcli avail # 4. List running jobs qzcli ls -c -r
bash# Interactive login qzcli login # With credentials qzcli login -u YOUR_USERNAME -p 'YOUR_PASSWORD' # Read password from stdin (for scripts) echo 'YOUR_PASSWORD' | qzcli login -u YOUR_USERNAME --password-stdin # Check current cookie qzcli cookie --show # Clear cookie qzcli cookie --clear
Note: qzcli avail auto-refreshes the cookie if it expires and credentials are configured.
bash# List cached workspaces qzcli res --list # Refresh all workspace resource cache (run this first!) qzcli res -u # Refresh a specific workspace qzcli res -w MY_WORKSPACE -u # Set a human-readable alias for a workspace qzcli res -w ws-xxxxxxxx --name "My Workspace"
bash# All workspaces qzcli avail # Including low-priority task nodes (slower but more accurate) qzcli avail --lp # Specific workspace qzcli avail -w MY_WORKSPACE # Find compute groups with N free nodes qzcli avail -n 4 # Export IDs for scripting qzcli avail -n 4 -e # Show idle node names qzcli avail -w MY_WORKSPACE -v
bash# Full interactive selection: workspace → project → compute group → spec qzcli create -i # Interactive for a specific workspace only qzcli create -i -w "My Workspace"
The TUI shows GPU type, availability, and spec status at each level. Press Enter/→ to go deeper, ← to go back.
bash# Using names (resolved from qzcli res cache) qzcli create \ --name "my-training-job" \ --command "bash /path/to/train.sh" \ --workspace "My Workspace" \ --compute-group "My Compute Group" \ --image YOUR_REGISTRY/team/image:tag \ --instances 4 \ --priority 10 # Using IDs directly qzcli create \ --name "my-job" \ --command "bash /path/to/train.sh" \ --workspace ws-YOUR_WORKSPACE_ID \ --compute-group lcg-YOUR_LCG_ID \ --spec YOUR_SPEC_ID \ --image YOUR_REGISTRY/team/image:tag \ --instances 4
Key parameters:
| Parameter | Default | Description | |-----------|---------|-------------| | --name / -n | required | Job name | | --command / -c | required | Command to run | | --workspace / -w | | Workspace name or ID (ws-...) | | --compute-group / -g | auto | Compute group name or ID (lcg-...) | | --spec / -s | auto | Resource spec ID | | --image / -m | | Docker image | | --instances | 1 | Number of instances | | --shm | 1200 | Shared memory (GiB) | | --priority | 10 | Priority (1–10) | | --dry-run | | Preview only, don't submit | | --json | | JSON output for scripting |
bash# Preview before submitting qzcli create --name test --command "echo hi" --workspace "My Workspace" \ --image YOUR_IMAGE --dry-run
bash# Pass vars directly — do NOT use "export VAR; bash script.sh" WORKSPACE_ID="ws-YOUR_WORKSPACE_ID" \ LCG_ID="lcg-YOUR_LCG_ID" \ SPEC_ID="YOUR_SPEC_ID" \ CHECKPOINT_DIR="/path/to/checkpoint" \ bash YOUR_SUBMIT_SCRIPT.sh
bashqzcli hpc \ --name "my-cpu-job" \ --workspace ws-YOUR_WORKSPACE_ID \ --compute-group lcg-YOUR_LCG_ID \ --predef-quota-id YOUR_QUOTA_ID \ --cpu 55 --mem-gi 300 --instances 30 \ --image YOUR_REGISTRY/team/cpu-image:tag \ --entrypoint "cd /path/to/dir && bash run.sh"
bash# Submit from config file qzcli batch batch_config.json --delay 3 # Preview all jobs qzcli batch batch_config.json --dry-run # Continue on error qzcli batch batch_config.json --continue-on-error
Config format (batch_config.json):
json{ "defaults": { "workspace": "ws-YOUR_WORKSPACE_ID", "compute_group": "lcg-YOUR_LCG_ID", "spec": "YOUR_SPEC_ID", "image": "YOUR_REGISTRY/team/image:tag", "instances": 4, "priority": 10 }, "matrix": { "checkpoint": ["/path/to/ckpt1", "/path/to/ckpt2"], "step": [50000, 100000] }, "name_template": "eval-{checkpoint_basename}-step{step}", "command_template": "bash eval.sh --checkpoint {checkpoint} --step {step}" }
Matrix keys are Cartesian-producted (2×2 = 4 jobs above). Use {key_basename} for path basenames.
bashfor step in 040000 050000 060000; do qzcli create \ --name "eval-step${step}" \ --command "bash eval.sh --step $step" \ --workspace "My Workspace" \ --compute-group "My Compute Group" \ --instances 4 sleep 3 done
bash# List jobs qzcli ls -c -w MY_WORKSPACE # specific workspace qzcli ls -c --all-ws # all workspaces qzcli ls -c -w MY_WORKSPACE -r # running only qzcli ls -c -w MY_WORKSPACE -n 50 # show 50 # Stop a job qzcli stop JOB_ID # Job status / details qzcli status JOB_ID # Watch all running jobs (refresh every 10s) qzcli watch -i 10 # Workspace view with GPU utilization qzcli ws qzcli ws -a # all projects qzcli ws -p "My Project"
| Problem | Cause | Fix | |---------|-------|-----| | Cookie expired | Session gap | Re-run qzcli login | | 未找到名称为 'xxx' 的工作空间 | Stale cache | Run qzcli res -u | | No resources in create -i | Cache empty | Run qzcli login && qzcli res -u | | qzcli-mcp not found | Not installed | cd qzcli_tool && pip install -e . | | Spec not in workspace | ID mismatch | Match spec ID to the correct workspace | | Silent job failure | Script sys.exit(0) | Check job logs directly | | zsh glob errors | Remote shell is zsh | Wrap commands in bash -c or use Python |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 7,403 | 2,799 | -62% | 1 | 1 | 0% | 1,340 | 2,889 | +116% | 0 | 0 | — |
case-02 | fail→pass | 16,668 | 2,367 | -86% | 1 | 1 | 0% | 1,327 | 2,857 | +115% | 0 | 0 | — |
case-03 | fail→pass | 9,179 | 3,680 | -60% | 1 | 1 | 0% | 1,902 | 3,245 | +71% | 0 | 0 | — |
case-04 | fail→pass | 12,306 | 4,152 | -66% | 1 | 1 | 0% | 2,262 | 3,207 | +42% | 0 | 0 | — |
case-05 | fail→pass | 9,170 | 1,442 | -84% | 1 | 1 | 0% | 1,633 | 2,658 | +63% | 0 | 0 | — |
case-06 | fail→pass | 6,634 | 2,241 | -66% | 1 | 1 | 0% | 1,209 | 2,730 | +126% | 0 | 0 | — |
case-07 | fail→pass | 6,752 | 1,159 | -83% | 1 | 1 | 0% | 1,189 | 2,577 | +117% | 0 | 0 | — |
case-08 | fail→pass | 6,159 | 1,810 | -71% | 1 | 1 | 0% | 954 | 2,648 | +178% | 0 | 0 | — |
case-09 | fail→pass | 4,349 | 1,609 | -63% | 1 | 1 | 0% | 828 | 2,645 | +219% | 0 | 0 | — |
case-15 | pass→pass | 11,687 | 1,703 | -85% | 1 | 1 | 0% | 2,367 | 2,683 | +13% | 0 | 0 | — |
case-10 | fail→pass | 6,907 | 2,938 | -57% | 1 | 1 | 0% | 1,549 | 3,069 | +98% | 0 | 0 | — |
case-11 | fail→pass | 4,825 | 1,839 | -62% | 1 | 1 | 0% | 871 | 2,676 | +207% | 0 | 0 | — |
case-12 | pass→pass | 7,468 | 4,033 | -46% | 1 | 1 | 0% | 1,389 | 3,199 | +130% | 0 | 0 | — |
case-13 | fail→pass | 10,360 | 2,934 | -72% | 1 | 1 | 0% | 2,147 | 2,993 | +39% | 0 | 0 | — |
case-14 | pass→pass | 9,987 | 1,713 | -83% | 1 | 1 | 0% | 1,763 | 2,717 | +54% | 0 | 0 | — |
case-16 | fail→pass | 5,350 | 1,875 | -65% | 1 | 1 | 0% | 1,006 | 2,710 | +169% | 0 | 0 | — |
case-17 | fail→pass | 7,805 | 1,755 | -78% | 1 | 1 | 0% | 1,324 | 2,667 | +101% | 0 | 0 | — |
case-18 | fail→pass | 8,268 | 1,928 | -77% | 1 | 1 | 0% | 1,559 | 2,679 | +72% | 0 | 0 | — |
case-19 | fail→pass | 7,086 | 15,098 | +113% | 1 | 1 | 0% | 1,335 | 2,736 | +105% | 0 | 0 | — |
case-20 | pass→pass | 13,070 | 10,078 | -23% | 1 | 1 | 0% | 2,719 | 4,452 | +64% | 0 | 0 | — |
case-21 | pass→pass | 11,019 | 8,037 | -27% | 1 | 1 | 0% | 2,328 | 4,028 | +73% | 0 | 0 | — |
case-22 | pass→pass | 8,301 | 7,048 | -15% | 1 | 1 | 0% | 1,856 | 4,044 | +118% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +73 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.