Login
Download
Skill UI
Browse and discover
15856+
curated skills
All
Development
Artificial Intelligence
Design & Creative
Product & Business
Data Science
Marketing
Soft Skills
Productivity
Engineering
Languages
Search
Consumer
, found
2
results
Default
Newest
Most Downloaded
GGUF Quantization for Efficient LLM Inference
gguf-quantization
Orchestra-Research/AI-Research-SKILLs
403
This guide details the use of the GGUF format and quantization techniques for running large language models (LLMs) efficiently. GGUF enables optimal inference on consumer hardware, Apple Silicon, and systems requiring CPU-only operation. By compressing models using various K-quant methods (like Q4_K_M), developers can significantly reduce memory footprint and hardware requirements while maintaining high performance.
View Details
GPTQ Quantization Guide
gptq
Orchestra-Research/AI-Research-SKILLs
81
GPTQ enables post-training 4-bit quantization for large LLMs, delivering up to 4× memory reduction and 3-4× faster inference with under 2% perplexity loss on consumer GPUs through AutoGPTQ, PEFT, and Transformer integrations.
View Details
1
Language
简体中文
English