Blog

PiRouter · Engineering Journal

Blog

Engineering notes on LLM routing, gateways, and AI infrastructure.

Dumbbell chart of bartowski's Nemotron-3.5-Lightning GGUF files: bits per weight the filename claims against bits per weight measured from the tensor table, every k- and i-quant rung landing at or above a 4.5 bpw floor

Latest ·

Why your IQ2 GGUF is the same size as IQ4: the 256-block fallback in llama-quantize

A tensor width not divisible by 256 makes llama-quantize fall back to ~4.5 bpw and keep the low-bit filename. We measured Flash-Next's 1.56 bpw file at 3.28.

Linden Kern

More from the journal

Matrix of GLM-5.3's sixteen OpenRouter endpoints on 2026-08-29 05:24 UTC, grouped by self-declared quantization: bf16 one host at $1.20/$4.00, fp8 seven hosts at $1.25–1.40 in, fp4 two hosts at $1.40/$4.40, unknown six hosts at $1.40/$4.40

GLM-5.3 open weights, day one: 16 providers, five prices, four quantizations — and BF16 is the cheapest

Linden Kern
Timeline of the SpaceX–Cursor deal and OpenAI's response: 21 April option announced, 14 August acquisition completed, 28 August notice given, 12 November proposed shutoff

Cursor loses OpenAI models on November 12: the change-of-control clause behind the date, and the five questions to ask whoever holds your key

Leo Kaka
Matrix of seven protocol dimensions across OpenAI Chat Completions, Anthropic Messages and Gemini generateContent, with the field name each vendor uses

One API for OpenAI, Anthropic and Google models: what translates, and the seven things that leak through

Leo Kaka
Four-stage diagram for Kimi K3: 2.8T weights free to download → host needs 8× GB300 or more → reported revenue share of up to 30% → token auditing unresolved. Below it, the three license thresholds: $20M over 12 months, 100M MAU or $20M a month, internal use and certified partners exempt

Kimi K3 license, commercial use: the $20M MaaS clause, the 100M-MAU badge, and the 30% Moonshot wants from Azure, AWS and Google

Linden Kern
Matrix of six LLM API products offered as OpenRouter alternatives — Cloudflare AI Gateway, AWS Bedrock, Google Vertex, Groq, LiteLLM and Portkey — against four production columns: cross-vendor failover, a spend cap below the account, the Anthropic Messages API, and a zero-data-retention control, each cell marked available, gated or absent according to the vendor's own documentation

OpenRouter alternatives for production workloads: what Cloudflare, Bedrock, Vertex, Groq, LiteLLM and Portkey actually document

Leo Kaka
Hand-drawn illustration: an open-air bazaar of model makers' stalls; at the gate a brass shovel-maker's cart has parked while a porter ties a single yellow reservation ribbon to the gatepost

Nvidia in talks to buy Hugging Face: the investor refused in January is the buyer in August

Linden Kern
Hand-drawn illustration: one open team wallet on a wooden workbench with five hands reaching in at once, a single yellow coin among grey tokens, and a brass padlock hanging unlatched from the clasp

Shared team wallet for LLM APIs: who actually has one, and what each cap misses

Leo Kaka
Matrix of failover trigger conditions across seven LLM gateways: 5xx, timeouts and 429 are covered everywhere, by default or opt-in; 404 is opt-in or unstated in most; the row for a 200 whose provider state changed is empty in every column, with Portkey's guardrail path as the one partial exception noted in the text

Automatic failover in LiteLLM, Portkey and OpenRouter: what actually triggers it, and the 429 no gateway tracks

Leo Kaka
Matrix of three LLM routers, LiteLLM, OpenRouter and Portkey, against four budget dimensions: the per-key cap parameter, the shared wallet, when the check runs, and the HTTP status returned when the cap is hit

Per-key spending limits: what LiteLLM, OpenRouter and Portkey enforce — and the four leaks a cap doesn't close

Leo Kaka
Path diagram: model tokens become Qwen3 Coder XML, enter the vLLM tool parser, reach a highlighted eval() call and then the host; below it a version bar showing eval() present from v0.10.0 through v0.10.1 and removed in v0.10.1.1

Is the vLLM tool-call parser safe? CVE-2025-9141, the 29-day eval() window, and four checks a gateway owes you

Leo Kaka
Timeline of two distribution-layer deals in 2026: Hugging Face declines Nvidia's $500M at a $7B valuation in January, OpenRouter announces it is joining Stripe on 19 August, Hugging Face is reported exploring a sale at $13B or more on 23 August

Hugging Face said no to one dominant investor. Eight months later it is exploring a sale

Linden Kern
Four price-change notice periods compared: OpenAI fourteen days, Anthropic thirty days or on notice, Google thirty days with new paid services immediate, DeepSeek no committed period

Price-change notice: 14 days, 30 days, or nothing at all

PiRouter Team
The five MCP roadmap priority areas listed in order, with agent identity and improved primitives marked as the two that land on the layer in front of the server

The MCP roadmap hands agent identity to your gateway

PiRouter Team
Bar chart of top-1 token flips against the BF16 reference at 88k context: INT8 about 15 percent, official FP8 about 20 percent highlighted, AWQ W4A16 about 30 percent, NVFP4 about 50 percent

FP8 flips 20% of top-1 tokens, and tool calls break first

Linden Kern
Unsloth's Qwen3.8-27B GGUF chart: top-1% accuracy against file size for four quantization providers

Qwen3.8-27B: seven endpoints, one name, two prices

Linden Kern
GitHub's own growth charts from the August 17 post-mortem: merged pull requests, commits and new repositories per month, 2023–2026 (The GitHub Blog)

Retry storms: 9K to 100K RPS, and four rules that stop them

Leo Kaka
Cover card: A price sheet, and a model with no name

ox-alpha vs deepseek-v4-flash: what a stealth model hides

Linden Kern
GLM-5.3 launch benchmark chart — one of the two time-of-day pricing moves this month rode in alongside these bars

DeepSeek peak pricing: which 30% of your bill can move

Leo Kaka
The source post-mortem's cover art: server racks trading cables in the dark (The Daily Brief)

Routers aren't dead — per-request routing is

Linden Kern
Cover: your provider's billing system is part of your failure domain

Tier 3 to Tier 1 mid-flight: the outage no 5xx will catch

Leo Kaka

Browse the journal

All 21 articles

Every article in one list, newest first, filterable by topic, author and month. Or go straight to a topic or an author.