Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Fine-tune LLMs with LlamaFactory — register datasets, train via YAML configs, merge LoRA adapters and serve the result.
.claude/skills/prism-shadow-llamafactory/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -38% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 3% | 0% |
LlamaFactory fine-tunes open-weight LLMs (LoRA/QLoRA and full-parameter; SFT, DPO and more) through the llamafactory-cli command driven by YAML configs.
If the user's message only invokes this skill (e.g. "use llamafactory skill") without a concrete request, ask the user what they want to fine-tune. Do not run any command until the goal is clear.
Confirm before training:
nvidia-smi) — it bounds the model size and method; LoRA needs far less than full fine-tuning.bashgit clone --depth 1 https://github.com/hiyouga/LlamaFactory.git cd LlamaFactory pip install -e . pip install -r requirements/metrics.txt # optional: evaluation metrics
Register every dataset in data/dataset_info.json; the alpaca and sharegpt formats are supported. A minimal local entry:
json"my_dataset": { "file_name": "my_dataset.json" }
alpaca rows carry instruction / input / output; sharegpt rows carry a conversations list. Put the data file under data/ next to the registry.
Training is driven by a YAML config. Start from the shipped example examples/train_lora/qwen3_lora_sft.yaml, or save a minimal config as my_sft.yaml, e.g. for Qwen/Qwen3-1.7B:
yamlmodel_name_or_path: Qwen/Qwen3-1.7B trust_remote_code: true stage: sft do_train: true finetuning_type: lora lora_rank: 8 lora_target: all dataset: my_dataset template: qwen3 output_dir: saves/qwen3-1.7b/lora/sft learning_rate: 1.0e-4 num_train_epochs: 3.0 bf16: true
bashllamafactory-cli train my_sft.yaml
llamafactory-cli webui launches the no-code web UI for the same workflow.
Merge the LoRA adapter into the base weights for standalone serving. Start from examples/merge_lora/qwen3_lora_sft.yaml, pointing model_name_or_path, adapter_name_or_path and template at your run (never merge into a quantized base):
yamlmodel_name_or_path: Qwen/Qwen3-1.7B adapter_name_or_path: saves/qwen3-1.7b/lora/sft template: qwen3 trust_remote_code: true export_dir: saves/qwen3-1.7b-sft-merged
bashllamafactory-cli export my_merge.yaml
Both commands take an inference config — derive it from examples/inference/qwen3_lora_sft.yaml, again pointing the model, adapter and template at your run:
yamlmodel_name_or_path: Qwen/Qwen3-1.7B adapter_name_or_path: saves/qwen3-1.7b/lora/sft template: qwen3 infer_backend: huggingface trust_remote_code: true
bashllamafactory-cli chat my_infer.yaml # interactive chat with the tuned model llamafactory-cli api my_infer.yaml # OpenAI-compatible API server
Serve the merged export as a standalone endpoint — vLLM serves the export directory directly, while Ollama needs an import first (a Modelfile with FROM /path/to/export, then ollama create; supported model architectures only) — then register the endpoint with PenguinHarness so agents can build, evaluate and tune AI apps on the fine-tuned model end to end.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,830 | 12,365 | -27% | 1 | 1 | 0% | 2,256 | 2,756 | +22% | 0 | 0 | — |
case-02 | fail→pass | 7,541 | 3,588 | -52% | 1 | 1 | 0% | 1,429 | 1,744 | +22% | 0 | 0 | — |
case-03 | fail→pass | 13,769 | 8,544 | -38% | 1 | 1 | 0% | 1,686 | 1,694 | +0% | 0 | 0 | — |
case-04 | fail→pass | 17,200 | 7,787 | -55% | 1 | 1 | 0% | 2,353 | 1,459 | -38% | 0 | 0 | — |
case-05 | pass→pass | 13,092 | 6,586 | -50% | 1 | 1 | 0% | 2,250 | 2,233 | -1% | 0 | 0 | — |
case-06 | fail→pass | 15,683 | 11,634 | -26% | 1 | 1 | 0% | 2,181 | 2,247 | +3% | 0 | 0 | — |
case-07 | pass→pass | 21,198 | 17,420 | -18% | 1 | 1 | 0% | 3,366 | 3,754 | +12% | 0 | 0 | — |
case-08 | pass→pass | 12,590 | 4,351 | -65% | 1 | 1 | 0% | 1,486 | 1,849 | +24% | 0 | 0 | — |
case-09 | pass→pass | 12,943 | 8,066 | -38% | 1 | 1 | 0% | 1,619 | 1,670 | +3% | 0 | 0 | — |
case-10 | pass→pass | 10,099 | 8,449 | -16% | 1 | 1 | 0% | 987 | 1,741 | +76% | 0 | 0 | — |
case-11 | pass→pass | 11,688 | 2,152 | -82% | 1 | 1 | 0% | 890 | 1,389 | +56% | 0 | 0 | — |
case-12 | pass→pass | 7,149 | 7,873 | +10% | 1 | 1 | 0% | 1,354 | 1,571 | +16% | 0 | 0 | — |
case-13 | pass→pass | 9,538 | 6,688 | -30% | 1 | 1 | 0% | 778 | 1,276 | +64% | 0 | 0 | — |
case-14 | fail→pass | 11,945 | 2,185 | -82% | 1 | 1 | 0% | 1,242 | 1,399 | +13% | 0 | 0 | — |
case-15 | pass→pass | 15,580 | 2,515 | -84% | 1 | 1 | 0% | 1,473 | 1,509 | +2% | 0 | 0 | — |
case-16 | pass→pass | 16,622 | 12,096 | -27% | 1 | 1 | 0% | 1,990 | 2,184 | +10% | 0 | 0 | — |
case-17 | pass→pass | 4,658 | 8,118 | +74% | 1 | 1 | 0% | 891 | 1,601 | +80% | 0 | 0 | — |
case-18 | pass→pass | 12,463 | 2,170 | -83% | 1 | 1 | 0% | 1,299 | 1,379 | +6% | 0 | 0 | — |
case-19 | pass→pass | 5,282 | 15,071 | +185% | 1 | 1 | 0% | 928 | 1,465 | +58% | 0 | 0 | — |
case-20 | pass→pass | 8,227 | 2,415 | -71% | 1 | 1 | 0% | 1,540 | 1,471 | -4% | 0 | 0 | — |
case-21 | pass→pass | 3,549 | 7,177 | +102% | 1 | 1 | 0% | 588 | 1,390 | +136% | 0 | 0 | — |
case-22 | fail→pass | 21,513 | 13,374 | -38% | 1 | 1 | 0% | 3,103 | 2,627 | -15% | 0 | 0 | — |
case-23 | pass→pass | 3,867 | 7,240 | +87% | 1 | 1 | 0% | 703 | 1,279 | +82% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +30 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.