GPTQ: Post-Training INT4
OBQ Foundation
GPTQ (Frantar et al., 2022) builds on Optimal Brain Quantization (OBQ), which itself extends the classic Optimal Brain Surgeon (OBS) framework. The central idea: when quantizing a weight, the quantization error can be partially compensated by adjusting the remaining unquantized weights in the same row.
OBQ formulates weight quantization as a constrained optimization. Given a weight matrix W ∈ ℝd_out × d_in and calibration inputs X, the objective is to find quantized weights Ŵ that minimize the layer-wise reconstruction error:
argmin_Ŵ ||WX - ŴX||₂²
OBQ quantizes weights one at a time, greedily choosing the weight whose quantization introduces the least additional error. After quantizing weight wij, it applies a closed-form update to all remaining weights in the same row to partially compensate the error.
Hessian-Based Weight Update
The Hessian of the layer-wise objective with respect to weights in row i is:
H = 2X·Xᵀ ∈ ℝd_in × d_in
This is the same Hessian for every row — it depends only on the input activations, not the weights. When we quantize weight wq at column q with quantization error δq = quant(wq) - wq, the optimal update to the remaining weights in the same row is:
δ_F = -δ_q · (HF,q / Hqq)
where F is the set of remaining (unquantized) column indices, HF,q is the column of the Hessian restricted to those indices, and Hqq is the diagonal entry. This is the OBS weight update — a rank-1 correction that minimizes the squared error increase.
The GPTQ Algorithm
GPTQ makes two critical departures from OBQ to achieve tractability on billion-parameter models:
- Fixed column ordering: Instead of greedily selecting the best column to quantize (O(d²) per step), GPTQ processes columns left-to-right. Empirically, the column order matters little because the Hessian-based compensation absorbs most of the error regardless of order.
- Block-wise processing: Columns are processed in blocks of B=128 columns. Within a block, updates are accumulated in registers. After the block completes, a single batched update is applied to all remaining columns. This converts O(d) memory writes per column into O(d/B) batched writes — a ~128× reduction in memory traffic.
The full algorithm processes each layer independently with the following steps:
def gptq_quantize_layer(W, X_cal, bits=4, group_size=128, block_size=128): """GPTQ: quantize one linear layer using calibration data.""" d_row, d_col = W.shape Q = torch.zeros_like(W) # quantized output H = 2 * (X_cal.T @ X_cal) # Hessian: d_col × d_col H += 1e-6 * torch.eye(d_col) # damping for numerical stability H_inv = torch.linalg.cholesky_inverse(torch.linalg.cholesky(H)) for block_start in range(0, d_col, block_size): block_end = min(block_start + block_size, d_col) W_block = W[:, block_start:block_end].clone() Err = torch.zeros_like(W_block) H_inv_block = H_inv[block_start:block_end, block_start:block_end] for j in range(block_end - block_start): col = block_start + j w = W_block[:, j] h_inv_jj = H_inv_block[j, j] # Quantize this column q = quantize_to_int(w, bits, group_size) Q[:, col] = q delta = (w - q) / h_inv_jj # scaled quantization error Err[:, j] = delta # Update remaining columns in this block W_block[:, j+1:] -= delta.unsqueeze(1) * H_inv_block[j, j+1:].unsqueeze(0) # Apply accumulated error to all remaining columns beyond this block W[:, block_end:] -= Err @ H_inv[block_start:block_end, block_end:] return Q
Group Quantization
Pure per-channel INT4 quantization uses one scale and zero-point per output channel (row). This works well for INT8, but at 4-bit precision the dynamic range per row is often too varied, causing large quantization errors for outlier columns.
Group quantization subdivides each row into groups of group_size consecutive weights (typically 128), each with its own scale and zero-point. This significantly improves accuracy at minimal storage overhead.
Per-Channel (group_size = d_in)
One scale per row. Overhead: 2 bytes per row. Quality degrades with wider layers. Perplexity penalty: +0.5–2.0 on WikiText-2 for LLaMA-7B at INT4.
Group-128 (group_size = 128)
One scale per 128 weights. Overhead: ~0.125 bits/weight extra (FP16 scale per 128 INT4 values = 16/(128×4) = 3.1% overhead). Perplexity penalty: +0.1–0.3 — nearly lossless.
The storage per weight with group quantization is: 4 bits (weight) + 16 bits / group_size (scale) + 16 bits / group_size (zero-point) = 4 + 32/128 = 4.25 bits/weight for group_size=128. Some implementations omit the zero-point for symmetric quantization, giving 4.125 bits/weight.
Memory Savings & Perplexity
GPTQ INT4 quantization dramatically reduces model size and memory requirements:
| Model | FP16 Size | INT4 GPTQ | Compression | PPL (FP16) | PPL (INT4) |
|---|---|---|---|---|---|
| LLaMA-7B | 13.5 GB | 3.6 GB | 3.75× | 5.68 | 5.85 |
| LLaMA-13B | 26 GB | 6.8 GB | 3.82× | 5.09 | 5.20 |
| LLaMA-30B | 62 GB | 16.2 GB | 3.83× | 4.10 | 4.20 |
| LLaMA-65B | 130 GB | 33.5 GB | 3.88× | 3.53 | 3.60 |
Calibration uses only 128 random samples from C4 (typically 2048 tokens each). The calibration set is used to compute the Hessian H = XXᵀ per layer. GPTQ is remarkably insensitive to the calibration data — using Wikipedia vs. C4 vs. code changes perplexity by <0.05 in most cases.
Running GPTQ with AutoGPTQ
AutoGPTQ provides a practical interface for GPTQ quantization with HuggingFace integration:
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig from transformers import AutoTokenizer import torch model_name = "meta-llama/Llama-2-70b-hf" quant_config = BaseQuantizeConfig( bits=4, # 4-bit quantization group_size=128, # group size for sub-channel quantization desc_act=False, # True = activation ordering (slower but better) damp_percent=0.01, # Hessian damping: H += damp * diag(H) sym=True, # symmetric quantization (no zero-point) model_seqlen=2048, # calibration sequence length ) # Load model in FP16 (needs ~140GB RAM or multi-GPU) tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoGPTQForCausalLM.from_pretrained( model_name, quant_config, torch_dtype=torch.float16, device_map="auto" ) # Prepare calibration data: 128 samples of 2048 tokens from datasets import load_dataset cal_data = load_dataset("allenai/c4", split="train", streaming=True) cal_texts = [sample["text"] for _, sample in zip(range(128), cal_data)] cal_tokens = tokenizer(cal_texts, return_tensors="pt", truncation=True, max_length=2048, padding=True) # Quantize: ~4 GPU-hours for 70B on single A100 model.quantize(cal_tokens["input_ids"]) # Save quantized model: ~35GB instead of ~140GB model.save_quantized("./llama-70b-gptq-int4") tokenizer.save_pretrained("./llama-70b-gptq-int4") # Load and run inference (fits on single 48GB GPU!) model = AutoGPTQForCausalLM.from_quantized( "./llama-70b-gptq-int4", device_map="auto", use_safetensors=True, inject_fused_attention=True, # fused kernels for faster inference ) output = model.generate(**tokenizer("The capital of France is", return_tensors="pt"))
Limitations & Trade-offs
Strengths
- Near-lossless INT4 for large models (>13B)
- 3.5–4× memory reduction
- Fast quantization (hours, not days)
- Excellent HuggingFace/vLLM ecosystem support
- Needs only 128 calibration samples
Weaknesses
- Requires full FP16 model in memory during quantization
- INT4 kernels are slower than FP16 on some hardware (INT4 dequant overhead)
- Small models (<7B) degrade more noticeably
- Not easily fine-tunable after quantization
- Activation ordering hurts inference speed
GPTQ remains the most widely deployed post-training quantization method for LLMs, forming the backbone of the Hugging Face GPTQ model ecosystem with thousands of pre-quantized models available. Its combination of strong quality preservation, fast quantization time, and broad tooling support makes it the default choice for deploying large models on constrained hardware.
Continue with worked examples
Continue with these focused notes: