Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or for faster inference (3-4× speedup) vs FP16. Integrates with transformers and PEFT for QLoRA fine-tuning.
.claude/skills/openlair-gptq/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 219% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 244% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 137% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 207% | 0% |
Post-training quantization method that compresses LLMs to 4-bit with minimal accuracy loss using group-wise quantization.
Use GPTQ when:
Use AWQ instead when:
Use bitsandbytes instead when:
bash# Install AutoGPTQ pip install auto-gptq # With Triton (Linux only, faster) pip install auto-gptq[triton] # With CUDA extensions (faster) pip install auto-gptq --no-build-isolation # Full installation pip install auto-gptq transformers accelerate
pythonfrom transformers import AutoTokenizer from auto_gptq import AutoGPTQForCausalLM # Load quantized model from HuggingFace model_name = "TheBloke/Llama-2-7B-Chat-GPTQ" model = AutoGPTQForCausalLM.from_quantized( model_name, device="cuda:0", use_triton=False # Set True on Linux for speed ) tokenizer = AutoTokenizer.from_pretrained(model_name) # Generate prompt = "Explain quantum computing" inputs = tokenizer(prompt, return_tensors="pt").to("cuda:0") outputs = model.generate(**inputs, max_new_tokens=200) print(tokenizer.decode(outputs[0]))
pythonfrom transformers import AutoTokenizer from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig from datasets import load_dataset # Load model model_name = "meta-llama/Llama-2-7b-chat-hf" tokenizer = AutoTokenizer.from_pretrained(model_name) # Quantization config quantize_config = BaseQuantizeConfig( bits=4, # 4-bit quantization group_size=128, # Group size (recommended: 128) desc_act=False, # Activation order (False for CUDA kernel) damp_percent=0.01 # Dampening factor ) # Load model for quantization model = AutoGPTQForCausalLM.from_pretrained( model_name, quantize_config=quantize_config ) # Prepare calibration data dataset = load_dataset("c4", split="train", streaming=True) calibration_data = [ tokenizer(example["text"])["input_ids"][:512] for example in dataset.take(128) ] # Quantize model.quantize(calibration_data) # Save quantized model model.save_quantized("llama-2-7b-gptq") tokenizer.save_pretrained("llama-2-7b-gptq") # Push to HuggingFace model.push_to_hub("username/llama-2-7b-gptq")
How GPTQ works:
Group size trade-off:
| Group Size | Model Size | Accuracy | Speed | Recommendation | |------------|------------|----------|-------|----------------| | -1 (per-column) | Smallest | Best | Slowest | Research only | | 32 | Smaller | Better | Slower | High accuracy needed | | 128 | Medium | Good | Fast | Recommended default | | 256 | Larger | Lower | Faster | Speed critical | | 1024 | Largest | Lowest | Fastest | Not recommended |
Example:
Weight matrix: [1024, 4096] = 4.2M elements
Group size = 128:
- Groups: 4.2M / 128 = 32,768 groups
- Each group: own 4-bit scale + zero-point
- Result: Better granularity → better accuracypythonfrom auto_gptq import BaseQuantizeConfig config = BaseQuantizeConfig( bits=4, # 4-bit quantization group_size=128, # Standard group size desc_act=False, # Faster CUDA kernel damp_percent=0.01 # Dampening factor )
Performance:
pythonconfig = BaseQuantizeConfig( bits=3, # 3-bit (more compression) group_size=128, # Keep standard group size desc_act=True, # Better accuracy (slower) damp_percent=0.01 )
Trade-off:
pythonconfig = BaseQuantizeConfig( bits=4, group_size=32, # Smaller groups (better accuracy) desc_act=True, # Activation reordering damp_percent=0.005 # Lower dampening )
Trade-off:
pythonmodel = AutoGPTQForCausalLM.from_quantized( model_name, device="cuda:0", use_exllama=True, # Use ExLlamaV2 exllama_config={"version": 2} )
Performance: 1.5-2× faster than Triton
python# Quantize with Marlin format config = BaseQuantizeConfig( bits=4, group_size=128, desc_act=False # Required for Marlin ) model.quantize(calibration_data, use_marlin=True) # Load with Marlin model = AutoGPTQForCausalLM.from_quantized( model_name, device="cuda:0", use_marlin=True # 2× faster on A100/H100 )
Requirements:
pythonmodel = AutoGPTQForCausalLM.from_quantized( model_name, device="cuda:0", use_triton=True # Linux only )
Performance: 1.2-1.5× faster than CUDA backend
pythonfrom transformers import AutoModelForCausalLM, AutoTokenizer # Load quantized model (transformers auto-detects GPTQ) model = AutoModelForCausalLM.from_pretrained( "TheBloke/Llama-2-13B-Chat-GPTQ", device_map="auto", trust_remote_code=False ) tokenizer = AutoTokenizer.from_pretrained("TheBloke/Llama-2-13B-Chat-GPTQ") # Use like any transformers model inputs = tokenizer("Hello", return_tensors="pt").to("cuda") outputs = model.generate(**inputs, max_new_tokens=100)
pythonfrom transformers import AutoModelForCausalLM from peft import prepare_model_for_kbit_training, LoraConfig, get_peft_model # Load GPTQ model model = AutoModelForCausalLM.from_pretrained( "TheBloke/Llama-2-7B-GPTQ", device_map="auto" ) # Prepare for LoRA training model = prepare_model_for_kbit_training(model) # LoRA config lora_config = LoraConfig( r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM" ) # Add LoRA adapters model = get_peft_model(model, lora_config) # Fine-tune (memory efficient!) # 70B model trainable on single A100 80GB
| Model | FP16 | GPTQ 4-bit | Reduction | |-------|------|------------|-----------| | Llama 2-7B | 14 GB | 3.5 GB | 4× | | Llama 2-13B | 26 GB | 6.5 GB | 4× | | Llama 2-70B | 140 GB | 35 GB | 4× | | Llama 3-405B | 810 GB | 203 GB | 4× |
Enables:
| Precision | Tokens/sec | vs FP16 | |-----------|------------|---------| | FP16 | 25 tok/s | 1× | | GPTQ 4-bit (CUDA) | 85 tok/s | 3.4× | | GPTQ 4-bit (ExLlama) | 105 tok/s | 4.2× | | GPTQ 4-bit (Marlin) | 120 tok/s | 4.8× |
| Model | FP16 | GPTQ 4-bit (g=128) | Degradation | |-------|------|---------------------|-------------| | Llama 2-7B | 5.47 | 5.55 | +1.5% | | Llama 2-13B | 4.88 | 4.95 | +1.4% | | Llama 2-70B | 3.32 | 3.38 | +1.8% |
Excellent quality preservation - less than 2% degradation!
python# Automatic device mapping model = AutoGPTQForCausalLM.from_quantized( "TheBloke/Llama-2-70B-GPTQ", device_map="auto", # Automatically split across GPUs max_memory={0: "40GB", 1: "40GB"} # Limit per GPU ) # Manual device mapping device_map = { "model.embed_tokens": 0, "model.layers.0-39": 0, # First 40 layers on GPU 0 "model.layers.40-79": 1, # Last 40 layers on GPU 1 "model.norm": 1, "lm_head": 1 } model = AutoGPTQForCausalLM.from_quantized( model_name, device_map=device_map )
python# Offload some layers to CPU (for very large models) model = AutoGPTQForCausalLM.from_quantized( "TheBloke/Llama-2-405B-GPTQ", device_map="auto", max_memory={ 0: "80GB", # GPU 0 1: "80GB", # GPU 1 2: "80GB", # GPU 2 "cpu": "200GB" # Offload overflow to CPU } )
python# Process multiple prompts efficiently prompts = [ "Explain AI", "Explain ML", "Explain DL" ] inputs = tokenizer(prompts, return_tensors="pt", padding=True).to("cuda") outputs = model.generate( **inputs, max_new_tokens=100, pad_token_id=tokenizer.eos_token_id ) for i, output in enumerate(outputs): print(f"Prompt {i}: {tokenizer.decode(output)}")
TheBloke on HuggingFace:
Search:
bash# Find GPTQ models on HuggingFace https://huggingface.co/models?library=gptq
Download:
pythonfrom auto_gptq import AutoGPTQForCausalLM # Automatically downloads from HuggingFace model = AutoGPTQForCausalLM.from_quantized( "TheBloke/Llama-2-70B-Chat-GPTQ", device="cuda:0" )
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 10,083 | 4,396 | -56% | 1 | 1 | 0% | 1,867 | 4,563 | +144% | 0 | 0 | — |
case-02 | pass→pass | 6,645 | 3,906 | -41% | 1 | 1 | 0% | 1,207 | 4,499 | +273% | 0 | 0 | — |
case-03 | fail→pass | 17,076 | 10,998 | -36% | 1 | 1 | 0% | 2,998 | 5,797 | +93% | 0 | 0 | — |
case-04 | fail→pass | 19,520 | 4,414 | -77% | 1 | 1 | 0% | 1,446 | 4,617 | +219% | 0 | 0 | — |
case-05 | pass→pass | 10,320 | 4,865 | -53% | 1 | 1 | 0% | 1,988 | 4,715 | +137% | 0 | 0 | — |
case-06 | pass→pass | 10,981 | 5,463 | -50% | 1 | 1 | 0% | 1,966 | 4,690 | +139% | 0 | 0 | — |
case-07 | pass→pass | 10,291 | 5,920 | -42% | 1 | 1 | 0% | 1,992 | 4,922 | +147% | 0 | 0 | — |
case-08 | fail→pass | 30,124 | 9,128 | -70% | 1 | 1 | 0% | 1,616 | 5,565 | +244% | 0 | 0 | — |
case-09 | pass→pass | 10,170 | 6,795 | -33% | 1 | 1 | 0% | 1,880 | 5,075 | +170% | 0 | 0 | — |
case-10 | pass→pass | 14,449 | 14,365 | -1% | 1 | 1 | 0% | 2,542 | 6,511 | +156% | 0 | 0 | — |
case-11 | fail→fail | 8,747 | 7,633 | -13% | 1 | 1 | 0% | 1,864 | 5,168 | +177% | 0 | 0 | — |
case-12 | fail→fail | 12,038 | 8,486 | -30% | 1 | 1 | 0% | 2,345 | 5,324 | +127% | 0 | 0 | — |
case-13 | pass→pass | 12,057 | 10,579 | -12% | 1 | 1 | 0% | 2,390 | 5,895 | +147% | 0 | 0 | — |
case-14 | pass→pass | 8,127 | 5,611 | -31% | 1 | 1 | 0% | 1,654 | 4,891 | +196% | 0 | 0 | — |
case-15 | pass→pass | 3,543 | 2,132 | -40% | 1 | 1 | 0% | 742 | 4,196 | +465% | 0 | 0 | — |
case-16 | pass→pass | 3,379 | 1,530 | -55% | 1 | 1 | 0% | 603 | 3,952 | +555% | 0 | 0 | — |
case-17 | fail→pass | 8,414 | 3,254 | -61% | 1 | 1 | 0% | 1,873 | 4,445 | +137% | 0 | 0 | — |
case-18 | fail→pass | 32,106 | 11,425 | -64% | 1 | 1 | 0% | 1,980 | 6,083 | +207% | 0 | 0 | — |
case-19 | pass→pass | 9,252 | 6,685 | -28% | 1 | 1 | 0% | 1,920 | 5,122 | +167% | 0 | 0 | — |
case-20 | pass→pass | 3,440 | 2,535 | -26% | 1 | 1 | 0% | 708 | 4,221 | +496% | 0 | 0 | — |
case-21 | pass→pass | 7,161 | 4,307 | -40% | 1 | 1 | 0% | 1,368 | 4,630 | +238% | 0 | 0 | — |
case-22 | pass→pass | 8,682 | 5,734 | -34% | 1 | 1 | 0% | 1,531 | 4,726 | +209% | 0 | 0 | — |
case-23 | pass→pass | 3,726 | 2,463 | -34% | 1 | 1 | 0% | 667 | 4,201 | +530% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +22 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.