Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when fine-tuning LLMs, training custom models, or adapting foundation models for specific tasks. Invoke for configuring LoRA/QLoRA adapters, preparing JSONL training datasets, setting hyperparameters for fine-tuning runs, adapter training, transfer learning, finetuning with Hugging Face PEFT, OpenAI fine-tuning, instruction tuning, RLHF, DPO, or quantizing and deploying fine-tuned models. Trigger terms include: LoRA, QLoRA, PEFT, finetuning, fine-tuning, adapter tuning, LLM training, model t
.claude/skills/jeffallan-fine-tuning-expert/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 97% | 66 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 323% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 12% | 0% |
Senior ML engineer specializing in LLM fine-tuning, parameter-efficient methods, and production model optimization.
python validate_dataset.py --input data.jsonl — fix all errors before proceedingLoad detailed guidance based on context:
| Topic | Reference | Load When | |-------|-----------|-----------| | LoRA/PEFT | references/lora-peft.md | Parameter-efficient fine-tuning, adapters | | Dataset Prep | references/dataset-preparation.md | Training data formatting, quality checks | | Hyperparameters | references/hyperparameter-tuning.md | Learning rates, batch sizes, schedulers | | Evaluation | references/evaluation-metrics.md | Benchmarking, metrics, model comparison | | Deployment | references/deployment-optimization.md | Model merging, quantization, serving |
pythonfrom datasets import load_dataset from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments from peft import LoraConfig, get_peft_model, TaskType from trl import SFTTrainer import torch # 1. Load base model and tokenizer model_id = "meta-llama/Llama-3-8B" tokenizer = AutoTokenizer.from_pretrained(model_id) tokenizer.pad_token = tokenizer.eos_token model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", ) # 2. Configure LoRA adapter lora_config = LoraConfig( task_type=TaskType.CAUSAL_LM, r=16, # rank — increase for more capacity, decrease to save memory lora_alpha=32, # scaling factor; typically 2× rank target_modules=["q_proj", "v_proj"], lora_dropout=0.05, bias="none", ) model = get_peft_model(model, lora_config) model.print_trainable_parameters() # verify: should be ~0.1–1% of total params # 3. Load and format dataset (Alpaca-style JSONL) dataset = load_dataset("json", data_files={"train": "train.jsonl", "test": "test.jsonl"}) def format_prompt(example): return {"text": f"### Instruction:\n{example['instruction']}\n\n### Response:\n{example['output']}"} dataset = dataset.map(format_prompt) # 4. Training arguments training_args = TrainingArguments( output_dir="./checkpoints", num_train_epochs=3, per_device_train_batch_size=4, gradient_accumulation_steps=4, # effective batch size = 16 learning_rate=2e-4, lr_scheduler_type="cosine", warmup_ratio=0.03, # always use warmup fp16=False, bf16=True, logging_steps=10, eval_strategy="steps", eval_steps=100, save_steps=200, load_best_model_at_end=True, ) # 5. Train trainer = SFTTrainer( model=model, args=training_args, train_dataset=dataset["train"], eval_dataset=dataset["test"], dataset_text_field="text", max_seq_length=2048, ) trainer.train() # 6. Save adapter weights only model.save_pretrained("./lora-adapter") tokenizer.save_pretrained("./lora-adapter")
QLoRA variant — add these lines before loading the model to enable 4-bit quantization:
pythonfrom transformers import BitsAndBytesConfig bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True, ) model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb_config, device_map="auto")
Merge adapter into base model for deployment:
pythonfrom peft import PeftModel base = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16) merged = PeftModel.from_pretrained(base, "./lora-adapter").merge_and_unload() merged.save_pretrained("./merged-model")
When implementing fine-tuning, always provide:
TrainingArguments + LoraConfig block, commented)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 29,837 | 26,143 | -12% | 1 | 1 | 0% | 6,239 | 7,801 | +25% | 0 | 0 | — |
case-02 | fail→fail | 24,275 | 25,128 | +4% | 1 | 1 | 0% | 6,229 | 7,791 | +25% | 0 | 0 | — |
case-03 | fail→pass | 27,006 | 27,343 | +1% | 1 | 1 | 0% | 6,224 | 7,786 | +25% | 0 | 0 | — |
case-04 | pass→pass | 20,315 | 23,694 | +17% | 1 | 1 | 0% | 4,103 | 4,594 | +12% | 0 | 0 | — |
case-05 | pass→pass | 19,992 | 20,190 | +1% | 1 | 1 | 0% | 3,148 | 5,321 | +69% | 0 | 0 | — |
case-06 | fail→pass | 44,756 | 26,706 | -40% | 1 | 1 | 0% | 1,720 | 7,283 | +323% | 0 | 0 | — |
case-07 | pass→pass | 10,338 | 17,205 | +66% | 1 | 1 | 0% | 2,374 | 5,655 | +138% | 0 | 0 | — |
case-08 | pass→pass | 8,631 | 5,163 | -40% | 1 | 1 | 0% | 1,669 | 2,692 | +61% | 0 | 0 | — |
case-09 | pass→pass | 8,110 | 7,905 | -3% | 1 | 1 | 0% | 1,678 | 3,237 | +93% | 0 | 0 | — |
case-10 | fail→fail | 5,847 | 7,653 | +31% | 1 | 1 | 0% | 1,232 | 2,903 | +136% | 0 | 0 | — |
case-11 | pass→pass | 8,712 | 7,919 | -9% | 1 | 1 | 0% | 1,847 | 3,237 | +75% | 0 | 0 | — |
case-12 | pass→pass | 11,923 | 16,489 | +38% | 1 | 1 | 0% | 2,374 | 4,840 | +104% | 0 | 0 | — |
case-13 | fail→pass | 6,977 | 1,332 | -81% | 1 | 1 | 0% | 1,192 | 1,778 | +49% | 0 | 0 | — |
case-14 | pass→pass | 8,315 | 7,981 | -4% | 1 | 1 | 0% | 1,673 | 3,148 | +88% | 0 | 0 | — |
case-15 | fail→pass | 14,267 | 16,184 | +13% | 1 | 1 | 0% | 3,130 | 5,265 | +68% | 0 | 0 | — |
case-16 | pass→pass | 11,411 | 9,619 | -16% | 1 | 1 | 0% | 2,123 | 3,357 | +58% | 0 | 0 | — |
case-17 | pass→pass | 3,731 | 10,633 | +185% | 1 | 1 | 0% | 692 | 2,448 | +254% | 0 | 0 | — |
case-18 | pass→pass | 9,232 | 8,511 | -8% | 1 | 1 | 0% | 1,993 | 3,352 | +68% | 0 | 0 | — |
case-19 | pass→pass | 8,311 | 7,064 | -15% | 1 | 1 | 0% | 1,524 | 2,824 | +85% | 0 | 0 | — |
case-20 | pass→pass | 4,405 | 4,128 | -6% | 1 | 1 | 0% | 819 | 2,221 | +171% | 0 | 0 | — |
case-21 | fail→fail | 16,668 | 17,039 | +2% | 1 | 1 | 0% | 3,261 | 4,729 | +45% | 0 | 0 | — |
case-22 | pass→pass | 7,573 | 5,557 | -27% | 1 | 1 | 0% | 1,501 | 2,716 | +81% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.