Topic
#quantization
4 articles

Why your IQ2 GGUF is the same size as IQ4: the 256-block fallback in llama-quantize
Linden Kern

GLM-5.3 open weights, day one: 16 providers, five prices, four quantizations — and BF16 is the cheapest
Linden Kern

FP8 flips 20% of top-1 tokens, and tool calls break first
Linden Kern

Qwen3.8-27B: seven endpoints, one name, two prices
Linden Kern