← All Posts

GPTQ: Post-Training INT4

OBQ Foundation

GPTQ (Frantar et al., 2022) builds on Optimal Brain Quantization (OBQ), which itself extends the classic Optimal Brain Surgeon (OBS) framework. The central idea: when quantizing a weight, the quantization error can be partially compensated by adjusting the remaining unquantized weights in the same row.

OBQ formulates weight quantization as a constrained optimization. Given a weight matrix W ∈ ℝd_out × d_in and calibration inputs X, the objective is to find quantized weights Ŵ that minimize the layer-wise reconstruction error:

argmin_Ŵ ||WX - ŴX||₂²

OBQ quantizes weights one at a time, greedily choosing the weight whose quantization introduces the least additional error. After quantizing weight wij, it applies a closed-form update to all remaining weights in the same row to partially compensate the error.

Key insight: Naive round-to-nearest quantization ignores weight correlations. By using second-order information (the Hessian), we can redistribute quantization error to preserve the layer's input-output mapping far more accurately.

Hessian-Based Weight Update

The Hessian of the layer-wise objective with respect to weights in row i is:

H = 2X·Xᵀ ∈ ℝd_in × d_in

This is the same Hessian for every row — it depends only on the input activations, not the weights. When we quantize weight wq at column q with quantization error δq = quant(wq) - wq, the optimal update to the remaining weights in the same row is:

δ_F = -δ_q · (HF,q / Hqq)

where F is the set of remaining (unquantized) column indices, HF,q is the column of the Hessian restricted to those indices, and Hqq is the diagonal entry. This is the OBS weight update — a rank-1 correction that minimizes the squared error increase.

Weight Matrix W (one row) Q Q w_q F F Hessian H = XXᵀ δ_q = quant(w_q) - w_q δ_F = -δ_q · H_Fq/H_qq Compensate remaining weights Updated Row Q F+δ F+δ Remaining weights adjusted to compensate Process: Left-to-right column scan → quantize column → update remaining → repeat GPTQ key trick: Fixed column order (no greedy search) + block processing for speed

The GPTQ Algorithm

GPTQ makes two critical departures from OBQ to achieve tractability on billion-parameter models:

  1. Fixed column ordering: Instead of greedily selecting the best column to quantize (O(d²) per step), GPTQ processes columns left-to-right. Empirically, the column order matters little because the Hessian-based compensation absorbs most of the error regardless of order.
  2. Block-wise processing: Columns are processed in blocks of B=128 columns. Within a block, updates are accumulated in registers. After the block completes, a single batched update is applied to all remaining columns. This converts O(d) memory writes per column into O(d/B) batched writes — a ~128× reduction in memory traffic.

The full algorithm processes each layer independently with the following steps:

def gptq_quantize_layer(W, X_cal, bits=4, group_size=128, block_size=128):
    """GPTQ: quantize one linear layer using calibration data."""
    d_row, d_col = W.shape
    Q = torch.zeros_like(W)                   # quantized output
    H = 2 * (X_cal.T @ X_cal)                 # Hessian: d_col × d_col
    H += 1e-6 * torch.eye(d_col)              # damping for numerical stability
    H_inv = torch.linalg.cholesky_inverse(torch.linalg.cholesky(H))

    for block_start in range(0, d_col, block_size):
        block_end = min(block_start + block_size, d_col)
        W_block = W[:, block_start:block_end].clone()
        Err = torch.zeros_like(W_block)
        H_inv_block = H_inv[block_start:block_end, block_start:block_end]

        for j in range(block_end - block_start):
            col = block_start + j
            w = W_block[:, j]
            h_inv_jj = H_inv_block[j, j]

            # Quantize this column
            q = quantize_to_int(w, bits, group_size)
            Q[:, col] = q
            delta = (w - q) / h_inv_jj       # scaled quantization error
            Err[:, j] = delta

            # Update remaining columns in this block
            W_block[:, j+1:] -= delta.unsqueeze(1) * H_inv_block[j, j+1:].unsqueeze(0)

        # Apply accumulated error to all remaining columns beyond this block
        W[:, block_end:] -= Err @ H_inv[block_start:block_end, block_end:]

    return Q
Complexity: GPTQ quantizes a 175B-parameter model in approximately 4 GPU-hours on a single A100, compared to OBQ which would take days. The block processing reduces the per-layer cost from O(d_col² · d_row) to O(d_col · d_row · B) memory operations.

Group Quantization

Pure per-channel INT4 quantization uses one scale and zero-point per output channel (row). This works well for INT8, but at 4-bit precision the dynamic range per row is often too varied, causing large quantization errors for outlier columns.

Group quantization subdivides each row into groups of group_size consecutive weights (typically 128), each with its own scale and zero-point. This significantly improves accuracy at minimal storage overhead.

Per-Channel (group_size = d_in)

One scale per row. Overhead: 2 bytes per row. Quality degrades with wider layers. Perplexity penalty: +0.5–2.0 on WikiText-2 for LLaMA-7B at INT4.

Group-128 (group_size = 128)

One scale per 128 weights. Overhead: ~0.125 bits/weight extra (FP16 scale per 128 INT4 values = 16/(128×4) = 3.1% overhead). Perplexity penalty: +0.1–0.3 — nearly lossless.

The storage per weight with group quantization is: 4 bits (weight) + 16 bits / group_size (scale) + 16 bits / group_size (zero-point) = 4 + 32/128 = 4.25 bits/weight for group_size=128. Some implementations omit the zero-point for symmetric quantization, giving 4.125 bits/weight.

Memory Savings & Perplexity

GPTQ INT4 quantization dramatically reduces model size and memory requirements:

Model FP16 Size INT4 GPTQ Compression PPL (FP16) PPL (INT4)
LLaMA-7B13.5 GB3.6 GB3.75×5.685.85
LLaMA-13B26 GB6.8 GB3.82×5.095.20
LLaMA-30B62 GB16.2 GB3.83×4.104.20
LLaMA-65B130 GB33.5 GB3.88×3.533.60
Scaling law for quantization: Larger models are more robust to quantization. The perplexity gap (INT4 vs FP16) decreases as model size increases. LLaMA-65B loses only 0.07 PPL at INT4, while LLaMA-7B loses 0.17. This makes GPTQ particularly effective for the largest models — exactly where memory savings matter most.

Calibration uses only 128 random samples from C4 (typically 2048 tokens each). The calibration set is used to compute the Hessian H = XXᵀ per layer. GPTQ is remarkably insensitive to the calibration data — using Wikipedia vs. C4 vs. code changes perplexity by <0.05 in most cases.

Running GPTQ with AutoGPTQ

AutoGPTQ provides a practical interface for GPTQ quantization with HuggingFace integration:

from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
from transformers import AutoTokenizer
import torch

model_name = "meta-llama/Llama-2-70b-hf"
quant_config = BaseQuantizeConfig(
    bits=4,                    # 4-bit quantization
    group_size=128,            # group size for sub-channel quantization
    desc_act=False,            # True = activation ordering (slower but better)
    damp_percent=0.01,         # Hessian damping: H += damp * diag(H)
    sym=True,                  # symmetric quantization (no zero-point)
    model_seqlen=2048,         # calibration sequence length
)

# Load model in FP16 (needs ~140GB RAM or multi-GPU)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoGPTQForCausalLM.from_pretrained(
    model_name,
    quant_config,
    torch_dtype=torch.float16,
    device_map="auto"
)

# Prepare calibration data: 128 samples of 2048 tokens
from datasets import load_dataset
cal_data = load_dataset("allenai/c4", split="train", streaming=True)
cal_texts = [sample["text"] for _, sample in zip(range(128), cal_data)]
cal_tokens = tokenizer(cal_texts, return_tensors="pt",
                        truncation=True, max_length=2048, padding=True)

# Quantize: ~4 GPU-hours for 70B on single A100
model.quantize(cal_tokens["input_ids"])

# Save quantized model: ~35GB instead of ~140GB
model.save_quantized("./llama-70b-gptq-int4")
tokenizer.save_pretrained("./llama-70b-gptq-int4")

# Load and run inference (fits on single 48GB GPU!)
model = AutoGPTQForCausalLM.from_quantized(
    "./llama-70b-gptq-int4",
    device_map="auto",
    use_safetensors=True,
    inject_fused_attention=True,  # fused kernels for faster inference
)
output = model.generate(**tokenizer("The capital of France is", return_tensors="pt"))
desc_act=True (activation ordering): Processes columns in order of decreasing activation magnitude rather than left-to-right. This yields ~0.05 lower perplexity but prevents efficient fused kernel execution because the column ordering is non-sequential, reducing inference throughput by ~15–20%. Recommended only when quality is paramount.

Limitations & Trade-offs

Strengths

  • Near-lossless INT4 for large models (>13B)
  • 3.5–4× memory reduction
  • Fast quantization (hours, not days)
  • Excellent HuggingFace/vLLM ecosystem support
  • Needs only 128 calibration samples

Weaknesses

  • Requires full FP16 model in memory during quantization
  • INT4 kernels are slower than FP16 on some hardware (INT4 dequant overhead)
  • Small models (<7B) degrade more noticeably
  • Not easily fine-tunable after quantization
  • Activation ordering hurts inference speed

GPTQ remains the most widely deployed post-training quantization method for LLMs, forming the backbone of the Hugging Face GPTQ model ecosystem with thousands of pre-quantized models available. Its combination of strong quality preservation, fast quantization time, and broad tooling support makes it the default choice for deploying large models on constrained hardware.

Continue with worked examples

Continue with these focused notes: