Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.
.claude/skills/openlair-quantizing-models-bitsandbytes/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 163% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 150% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 91% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 107% | 0% |
bitsandbytes reduces LLM memory by 50% (8-bit) or 75% (4-bit) with <1% accuracy loss.
Installation:
bashpip install bitsandbytes transformers accelerate
8-bit quantization (50% memory reduction):
pythonfrom transformers import AutoModelForCausalLM, BitsAndBytesConfig config = BitsAndBytesConfig(load_in_8bit=True) model = AutoModelForCausalLM.from_pretrained( "meta-llama/Llama-2-7b-hf", quantization_config=config, device_map="auto" ) # Memory: 14GB → 7GB
4-bit quantization (75% memory reduction):
pythonconfig = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16 ) model = AutoModelForCausalLM.from_pretrained( "meta-llama/Llama-2-7b-hf", quantization_config=config, device_map="auto" ) # Memory: 14GB → 3.5GB
Copy this checklist:
Quantization Loading:
- [ ] Step 1: Calculate memory requirements
- [ ] Step 2: Choose quantization level (4-bit or 8-bit)
- [ ] Step 3: Configure quantization
- [ ] Step 4: Load and verify modelStep 1: Calculate memory requirements
Estimate model memory:
FP16 memory (GB) = Parameters × 2 bytes / 1e9
INT8 memory (GB) = Parameters × 1 byte / 1e9
INT4 memory (GB) = Parameters × 0.5 bytes / 1e9
Example (Llama 2 7B):
FP16: 7B × 2 / 1e9 = 14 GB
INT8: 7B × 1 / 1e9 = 7 GB
INT4: 7B × 0.5 / 1e9 = 3.5 GBStep 2: Choose quantization level
| GPU VRAM | Model Size | Recommended | |----------|------------|-------------| | 8 GB | 3B | 4-bit | | 12 GB | 7B | 4-bit | | 16 GB | 7B | 8-bit or 4-bit | | 24 GB | 13B | 8-bit or 70B 4-bit | | 40+ GB | 70B | 8-bit |
Step 3: Configure quantization
For 8-bit (better accuracy):
pythonfrom transformers import BitsAndBytesConfig import torch config = BitsAndBytesConfig( load_in_8bit=True, llm_int8_threshold=6.0, # Outlier threshold llm_int8_has_fp16_weight=False )
For 4-bit (maximum memory savings):
pythonconfig = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16, # Compute in FP16 bnb_4bit_quant_type="nf4", # NormalFloat4 (recommended) bnb_4bit_use_double_quant=True # Nested quantization )
Step 4: Load and verify model
pythonfrom transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained( "meta-llama/Llama-2-13b-hf", quantization_config=config, device_map="auto", # Automatic device placement torch_dtype=torch.float16 ) tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-13b-hf") # Test inference inputs = tokenizer("Hello, how are you?", return_tensors="pt").to("cuda") outputs = model.generate(**inputs, max_length=50) print(tokenizer.decode(outputs[0])) # Check memory import torch print(f"Memory allocated: {torch.cuda.memory_allocated()/1e9:.2f}GB")
QLoRA enables fine-tuning large models on consumer GPUs.
Copy this checklist:
QLoRA Fine-tuning:
- [ ] Step 1: Install dependencies
- [ ] Step 2: Configure 4-bit base model
- [ ] Step 3: Add LoRA adapters
- [ ] Step 4: Train with standard TrainerStep 1: Install dependencies
bashpip install bitsandbytes transformers peft accelerate datasets
Step 2: Configure 4-bit base model
pythonfrom transformers import AutoModelForCausalLM, BitsAndBytesConfig import torch bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16, bnb_4bit_quant_type="nf4", bnb_4bit_use_double_quant=True ) model = AutoModelForCausalLM.from_pretrained( "meta-llama/Llama-2-7b-hf", quantization_config=bnb_config, device_map="auto" )
Step 3: Add LoRA adapters
pythonfrom peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training # Prepare model for training model = prepare_model_for_kbit_training(model) # Configure LoRA lora_config = LoraConfig( r=16, # LoRA rank lora_alpha=32, # LoRA alpha target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM" ) # Add LoRA adapters model = get_peft_model(model, lora_config) model.print_trainable_parameters() # Output: trainable params: 4.2M || all params: 6.7B || trainable%: 0.06%
Step 4: Train with standard Trainer
pythonfrom transformers import Trainer, TrainingArguments training_args = TrainingArguments( output_dir="./qlora-output", per_device_train_batch_size=4, gradient_accumulation_steps=4, num_train_epochs=3, learning_rate=2e-4, fp16=True, logging_steps=10, save_strategy="epoch" ) trainer = Trainer( model=model, args=training_args, train_dataset=train_dataset, tokenizer=tokenizer ) trainer.train() # Save LoRA adapters (only ~20MB) model.save_pretrained("./qlora-adapters")
Use 8-bit Adam/AdamW to reduce optimizer memory by 75%.
8-bit Optimizer Setup:
- [ ] Step 1: Replace standard optimizer
- [ ] Step 2: Configure training
- [ ] Step 3: Monitor memory savingsStep 1: Replace standard optimizer
pythonimport bitsandbytes as bnb from transformers import Trainer, TrainingArguments # Instead of torch.optim.AdamW model = AutoModelForCausalLM.from_pretrained("model-name") training_args = TrainingArguments( output_dir="./output", per_device_train_batch_size=8, optim="paged_adamw_8bit", # 8-bit optimizer learning_rate=5e-5 ) trainer = Trainer( model=model, args=training_args, train_dataset=train_dataset ) trainer.train()
Manual optimizer usage:
pythonimport bitsandbytes as bnb optimizer = bnb.optim.AdamW8bit( model.parameters(), lr=1e-4, betas=(0.9, 0.999), eps=1e-8 ) # Training loop for batch in dataloader: loss = model(**batch).loss loss.backward() optimizer.step() optimizer.zero_grad()
Step 2: Configure training
Compare memory:
Standard AdamW optimizer memory = model_params × 8 bytes (states)
8-bit AdamW memory = model_params × 2 bytes
Savings = 75% optimizer memory
Example (Llama 2 7B):
Standard: 7B × 8 = 56 GB
8-bit: 7B × 2 = 14 GB
Savings: 42 GBStep 3: Monitor memory savings
pythonimport torch before = torch.cuda.memory_allocated() # Training step optimizer.step() after = torch.cuda.memory_allocated() print(f"Memory used: {(after-before)/1e9:.2f}GB")
Use bitsandbytes when:
Use alternatives instead:
Issue: CUDA error during loading
Install matching CUDA version:
bash# Check CUDA version nvcc --version # Install matching bitsandbytes pip install bitsandbytes --no-cache-dir
Issue: Model loading slow
Use CPU offload for large models:
pythonmodel = AutoModelForCausalLM.from_pretrained( "model-name", quantization_config=config, device_map="auto", max_memory={0: "20GB", "cpu": "30GB"} # Offload to CPU )
Issue: Lower accuracy than expected
Try 8-bit instead of 4-bit:
pythonconfig = BitsAndBytesConfig(load_in_8bit=True) # 8-bit has <0.5% accuracy loss vs 1-2% for 4-bit
Or use NF4 with double quantization:
pythonconfig = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", # Better than fp4 bnb_4bit_use_double_quant=True # Extra accuracy )
Issue: OOM even with 4-bit
Enable CPU offload:
pythonmodel = AutoModelForCausalLM.from_pretrained( "model-name", quantization_config=config, device_map="auto", offload_folder="offload", # Disk offload offload_state_dict=True )
QLoRA training guide: See references/qlora-training.md for complete fine-tuning workflows, hyperparameter tuning, and multi-GPU training.
Quantization formats: See references/quantization-formats.md for INT8, NF4, FP4 comparison, double quantization, and custom quantization configs.
Memory optimization: See references/memory-optimization.md for CPU offloading strategies, gradient checkpointing, and memory profiling.
Supported platforms: NVIDIA GPUs (primary), AMD ROCm, Intel GPUs (experimental)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 6,875 | 2,625 | -62% | 1 | 1 | 0% | 1,423 | 3,740 | +163% | 0 | 0 | — |
case-02 | pass→pass | 7,035 | 3,339 | -53% | 1 | 1 | 0% | 1,532 | 3,834 | +150% | 0 | 0 | — |
case-03 | pass→pass | 11,359 | 6,032 | -47% | 1 | 1 | 0% | 2,300 | 4,387 | +91% | 0 | 0 | — |
case-04 | pass→pass | 11,092 | 7,092 | -36% | 1 | 1 | 0% | 2,201 | 4,558 | +107% | 0 | 0 | — |
case-05 | pass→pass | 3,181 | 3,285 | +3% | 1 | 1 | 0% | 694 | 3,859 | +456% | 0 | 0 | — |
case-06 | pass→pass | 3,063 | 2,765 | -10% | 1 | 1 | 0% | 686 | 3,703 | +440% | 0 | 0 | — |
case-07 | pass→pass | 2,940 | 2,303 | -22% | 1 | 1 | 0% | 630 | 3,621 | +475% | 0 | 0 | — |
case-08 | pass→pass | 3,581 | 2,390 | -33% | 1 | 1 | 0% | 686 | 3,632 | +429% | 0 | 0 | — |
case-13 | pass→pass | 8,931 | 4,178 | -53% | 1 | 1 | 0% | 1,718 | 3,976 | +131% | 0 | 0 | — |
case-09 | pass→pass | 3,665 | 3,719 | +1% | 1 | 1 | 0% | 778 | 3,940 | +406% | 0 | 0 | — |
case-10 | pass→pass | 4,460 | 3,223 | -28% | 1 | 1 | 0% | 911 | 3,787 | +316% | 0 | 0 | — |
case-11 | pass→pass | 5,265 | 3,157 | -40% | 1 | 1 | 0% | 1,070 | 3,785 | +254% | 0 | 0 | — |
case-12 | pass→pass | 3,437 | 4,449 | +29% | 1 | 1 | 0% | 618 | 4,102 | +564% | 0 | 0 | — |
case-14 | pass→pass | 5,696 | 4,157 | -27% | 1 | 1 | 0% | 1,092 | 3,980 | +264% | 0 | 0 | — |
case-15 | pass→pass | 7,963 | 2,499 | -69% | 1 | 1 | 0% | 1,488 | 3,628 | +144% | 0 | 0 | — |
case-16 | pass→pass | 14,171 | 6,082 | -57% | 1 | 1 | 0% | 2,669 | 4,278 | +60% | 0 | 0 | — |
case-17 | fail→pass | 9,819 | 1,608 | -84% | 1 | 1 | 0% | 1,689 | 3,408 | +102% | 0 | 0 | — |
case-18 | pass→pass | 12,007 | 7,024 | -42% | 1 | 1 | 0% | 2,368 | 4,413 | +86% | 0 | 0 | — |
case-19 | pass→pass | 10,371 | 9,064 | -13% | 1 | 1 | 0% | 1,956 | 4,964 | +154% | 0 | 0 | — |
case-20 | pass→pass | 15,996 | 11,727 | -27% | 1 | 1 | 0% | 2,942 | 5,441 | +85% | 0 | 0 | — |
case-21 | pass→pass | 10,931 | 9,955 | -9% | 1 | 1 | 0% | 2,043 | 5,036 | +147% | 0 | 0 | — |
case-22 | pass→pass | 7,456 | 6,887 | -8% | 1 | 1 | 0% | 1,532 | 4,443 | +190% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.