Reference-grade guide to shrinking and speeding up LLMs without retraining from scratch — numeric formats (FP8/INT8/INT4), PTQ methods (GPTQ, AWQ, SmoothQuant, bitsandbytes NF4, GGUF k-quants), KV-cache quantization, speculative decoding (Medusa/EAGLE/n-gram), and distillation — with concrete number
Reference-grade guide to shrinking and speeding up LLMs without retraining from scratch — numeric formats (FP8/INT8/INT4), PTQ methods (GPTQ, AWQ, SmoothQuant, bitsandbytes NF4, GGUF k-quants), KV-cache quantization, speculative decoding (Medusa/EAGLE/n-gram), and distillation — with concrete numbers, when each fits, and the quality cliffs.
Questa voce non pubblica alcun pacchetto npm, quindi Forge non ha un albero delle dipendenze per essa. È una lacuna di copertura, non l'affermazione che non abbia dipendenze.