<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>PiRouter Blog</title><description>Engineering notes on LLM routing, gateways, and AI infrastructure.</description><link>https://pirouter.ai/</link><language>en</language><item><title>Why your IQ2 GGUF is the same size as IQ4: the 256-block fallback in llama-quantize</title><link>https://pirouter.ai/blog/gguf-filename-is-a-request/</link><guid isPermaLink="true">https://pirouter.ai/blog/gguf-filename-is-a-request/</guid><description>A tensor width not divisible by 256 makes llama-quantize fall back to ~4.5 bpw and keep the low-bit filename. We measured Flash-Next&apos;s 1.56 bpw file at 3.28.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>quantization</category><category>models</category><category>local-llm</category><author>Linden Kern</author></item><item><title>GLM-5.3 open weights, day one: 16 providers, five prices, four quantizations — and BF16 is the cheapest</title><link>https://pirouter.ai/blog/glm-5-3-open-weights-day-one/</link><guid isPermaLink="true">https://pirouter.ai/blog/glm-5-3-open-weights-day-one/</guid><description>Fourteen hours after the weights landed, sixteen hosts served z-ai/glm-5.3 at five input prices and four declared precisions; the cheapest runs BF16.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>models</category><category>pricing</category><category>providers</category><category>quantization</category><author>Linden Kern</author></item><item><title>Cursor loses OpenAI models on November 12: the change-of-control clause behind the date, and the five questions to ask whoever holds your key</title><link>https://pirouter.ai/blog/model-access-change-of-control/</link><guid isPermaLink="true">https://pirouter.ai/blog/model-access-change-of-control/</guid><description>OpenAI gave Cursor 76 days under a change-of-control clause and excluded future models. What the clause means when your models arrive via someone else&apos;s key.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>providers</category><category>routing</category><category>failover</category><author>Leo Kaka</author></item><item><title>One API for OpenAI, Anthropic and Google models: what translates, and the seven things that leak through</title><link>https://pirouter.ai/blog/one-api-openai-anthropic-google/</link><guid isPermaLink="true">https://pirouter.ai/blog/one-api-openai-anthropic-google/</guid><description>Three ways to call OpenAI, Anthropic and Google through one API — a vendor&apos;s OpenAI-compatible endpoint, a self-hosted proxy, a hosted gateway — read against the official docs. All three flatten the request shape. None can flatten tool_choice, reasoning state, cache control, stream recovery or error codes.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>gateway</category><category>routing</category><category>agents</category><author>Leo Kaka</author></item><item><title>Kimi K3 license, commercial use: the $20M MaaS clause, the 100M-MAU badge, and the 30% Moonshot wants from Azure, AWS and Google</title><link>https://pirouter.ai/blog/open-weights-still-need-a-host/</link><guid isPermaLink="true">https://pirouter.ai/blog/open-weights-still-need-a-host/</guid><description>K3&apos;s weights are free; serving starts at 8× GB300, so a host stays. Reuters: Moonshot wants up to 30% from Azure, AWS and Google — token auditing is unresolved.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>providers</category><category>models</category><category>ecosystem</category><author>Linden Kern</author></item><item><title>OpenRouter alternatives for production workloads: what Cloudflare, Bedrock, Vertex, Groq, LiteLLM and Portkey actually document</title><link>https://pirouter.ai/blog/openrouter-alternatives-production/</link><guid isPermaLink="true">https://pirouter.ai/blog/openrouter-alternatives-production/</guid><description>Seven vendors&apos; docs, read 2026-08-29, on six production columns: failover triggers, spend caps, protocols, ZDR, rate-limit units, per-request logs. Every gap marked.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>gateway</category><category>providers</category><category>failover</category><author>Leo Kaka</author></item><item><title>Nvidia in talks to buy Hugging Face: the investor refused in January is the buyer in August</title><link>https://pirouter.ai/blog/nvidia-hf-talks/</link><guid isPermaLink="true">https://pirouter.ai/blog/nvidia-hf-talks/</guid><description>Business Insider reports Nvidia has been in talks to acquire Hugging Face at more than $13 billion — no deal yet, and the talks could still fall apart. In January, Hugging Face turned down $500M from the same company because it did not want a single dominant investor. Here is the timeline with the qualifiers kept intact, what the HN headline got wrong, and how the three neutrality measurements we published last week port from routers to model hubs.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><category>providers</category><category>ecosystem</category><category>opinion</category><author>Linden Kern</author></item><item><title>Shared team wallet for LLM APIs: who actually has one, and what each cap misses</title><link>https://pirouter.ai/blog/shared-team-wallet-llm-api/</link><guid isPermaLink="true">https://pirouter.ai/blog/shared-team-wallet-llm-api/</guid><description>People keep asking which LLM API aggregator has a shared team wallet. Read from the docs: OpenRouter and Vercel AI Gateway have a real shared balance; LiteLLM and Portkey have fine-grained member controls but no wallet. OpenRouter comes closest to both halves — until you hit the 10-member cap and the Enterprise gates. Here is the matrix, the fees, the reset cycles, and the caveats that don&apos;t appear on any pricing page.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate><category>cost</category><category>gateway</category><category>api-keys</category><author>Leo Kaka</author></item><item><title>Automatic failover in LiteLLM, Portkey and OpenRouter: what actually triggers it, and the 429 no gateway tracks</title><link>https://pirouter.ai/blog/failover-signals/</link><guid isPermaLink="true">https://pirouter.ai/blog/failover-signals/</guid><description>Seven LLM gateways&apos; failover docs, read verbatim: the trigger lists agree, every cooldown is a timer, and no routing policy records a 429 lasting eight hours.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate><category>reliability</category><category>failover</category><category>gateway</category><author>Leo Kaka</author></item><item><title>Per-key spending limits: what LiteLLM, OpenRouter and Portkey enforce — and the four leaks a cap doesn&apos;t close</title><link>https://pirouter.ai/blog/per-key-spending-limits/</link><guid isPermaLink="true">https://pirouter.ai/blog/per-key-spending-limits/</guid><description>How LiteLLM, OpenRouter and Portkey cap spend per API key — layer, timing, status code, reset — and the four leaks a cap doesn&apos;t close, read from the docs.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate><category>cost</category><category>gateway</category><category>api-keys</category><author>Leo Kaka</author></item><item><title>Is the vLLM tool-call parser safe? CVE-2025-9141, the 29-day eval() window, and four checks a gateway owes you</title><link>https://pirouter.ai/blog/tool-call-parser-attack-surface/</link><guid isPermaLink="true">https://pirouter.ai/blog/tool-call-parser-attack-surface/</guid><description>vLLM&apos;s Qwen3 Coder tool parser ran eval() on model output for 29 days in 2025 (CVE-2025-9141). Today: 0 eval(), 4 literal_eval. What a gateway can and cannot check.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate><category>security</category><category>gateway</category><category>local-llm</category><author>Leo Kaka</author></item><item><title>Hugging Face said no to one dominant investor. Eight months later it is exploring a sale</title><link>https://pirouter.ai/blog/distribution-layer-for-sale/</link><guid isPermaLink="true">https://pirouter.ai/blog/distribution-layer-for-sale/</guid><description>In January, Hugging Face turned down $500M from Nvidia because it did not want a single dominant investor. In August, it is exploring a sale at $13B or more — the same month OpenRouter announced it is joining Stripe. Both companies sell neutrality. Here are three things you can measure yourself to find out whether that promise survives an acquisition.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate><category>routing</category><category>providers</category><category>opinion</category><author>Linden Kern</author></item><item><title>Price-change notice: 14 days, 30 days, or nothing at all</title><link>https://pirouter.ai/blog/price-notice-is-a-contract-clause/</link><guid isPermaLink="true">https://pirouter.ai/blog/price-notice-is-a-contract-clause/</guid><description>Fourteen days, thirty days, thirty days, and none at all. In one week this August, one provider cancelled an increase it had already announced and another rebuilt its whole price sheet on three days&apos; notice — and both acted entirely within their own terms. Here are the four clauses side by side, where each one is buried, and what the floor means for how often you should re-check a price.</description><pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate><category>pricing</category><category>providers</category><category>cost</category><author>PiRouter Team</author></item><item><title>The MCP roadmap hands agent identity to your gateway</title><link>https://pirouter.ai/blog/mcp-roadmap-gateway/</link><guid isPermaLink="true">https://pirouter.ai/blog/mcp-roadmap-gateway/</guid><description>Two of the five priority areas in the MCP roadmap published on August 22 — agent identity and tool-list bloat — land in front of the server, not inside it. DPoP, Workload Identity Federation, ID-JAG and token exchange each move work to the layer you already run. What the 2026-07-28 stateless release already gave that layer, and which parts of the overreach critique are correct.</description><pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate><category>mcp</category><category>gateway</category><category>agents</category><author>PiRouter Team</author></item><item><title>FP8 flips 20% of top-1 tokens, and tool calls break first</title><link>https://pirouter.ai/blog/quantization-is-a-product-spec/</link><guid isPermaLink="true">https://pirouter.ai/blog/quantization-is-a-product-spec/</guid><description>A week of teacher-forced logit capture on one set of weights: at 88k context, official FP8 flips about 20% of top-1 tokens against BF16, NVFP4 about 50%. The interesting part is not the ladder. It is that the damage lands on tool calls — wrong interface, corrupted port number, a parameter name quietly truncated — and a routing table keyed on the model name cannot see any of the variables that caused it.</description><pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate><category>quantization</category><category>models</category><category>routing</category><author>Linden Kern</author></item><item><title>Qwen3.8-27B: seven endpoints, one name, two prices</title><link>https://pirouter.ai/blog/one-weight-many-prices/</link><guid isPermaLink="true">https://pirouter.ai/blog/one-weight-many-prices/</guid><description>One Apache-2.0 checkpoint reaches the market as two dozen GGUF files, seven API endpoints whose context windows run from 65,500 to 1,000,000 tokens, and — following one dated build across two hosts — two prices for a file that has not changed. The product is the tuple (weights, quantization, context window, version); a router keyed on the name is routing blind.</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate><category>models</category><category>quantization</category><category>routing</category><author>Linden Kern</author></item><item><title>Retry storms: 9K to 100K RPS, and four rules that stop them</title><link>https://pirouter.ai/blog/retry-storms/</link><guid isPermaLink="true">https://pirouter.ai/blog/retry-storms/</guid><description>GitHub&apos;s August 17 outage started with an autoscaling policy and spent its last four and a half hours fighting its own retries — one token service went from 7–9K to 70–100K requests per second on retries. The same arithmetic is sitting in your LLM SDK defaults and in your gateway&apos;s failover path. Here is the mechanism, the week&apos;s LLM-provider samples of it, and the four rules a gateway owes the providers behind it.</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate><category>reliability</category><category>failover</category><category>gateway</category><author>Leo Kaka</author></item><item><title>ox-alpha vs deepseek-v4-flash: what a stealth model hides</title><link>https://pirouter.ai/blog/stealth-models-and-price-sheets/</link><guid isPermaLink="true">https://pirouter.ai/blog/stealth-models-and-price-sheets/</guid><description>stealth/ox-alpha launched with a 1M context window, $0 pricing and no vendor name; deepseek-v4-flash-vision-exp launched the same day with nine benchmarks and a full price sheet. Both work. We compare the two listings on the four fields a router needs before either can enter a production routing table: identity confidence, retention policy, free-window expiry, and price-sheet completeness.</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate><category>models</category><category>routing</category><category>stealth-models</category><category>deepseek</category><author>Linden Kern</author></item><item><title>DeepSeek peak pricing: which 30% of your bill can move</title><link>https://pirouter.ai/blog/time-as-a-pricing-variable/</link><guid isPermaLink="true">https://pirouter.ai/blog/time-as-a-pricing-variable/</guid><description>DeepSeek now prices every line twice, peak and off-peak; GLM-5.3 launched with off-peak discounts. If 30% of your tokens are movable, a half-price window is worth up to 15% of the bill. The official rate cards, the peak windows converted to your timezone, and where the deadline-not-start-time decision belongs.</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate><category>pricing</category><category>routing</category><category>cost</category><author>Leo Kaka</author></item><item><title>Routers aren&apos;t dead — per-request routing is</title><link>https://pirouter.ai/blog/routers-arent-dead/</link><guid isPermaLink="true">https://pirouter.ai/blog/routers-arent-dead/</guid><description>A widely shared post-mortem says prompt caching killed the per-request model router, and its arithmetic holds up: a 10x cache discount beats a 2.5x list-price spread every time. We build a router. Here is what we co-sign, what we concede, and where routing actually earns its keep.</description><pubDate>Thu, 20 Aug 2026 00:00:00 GMT</pubDate><category>routing</category><category>caching</category><category>opinion</category><author>Linden Kern</author></item><item><title>Tier 3 to Tier 1 mid-flight: the outage no 5xx will catch</title><link>https://pirouter.ai/blog/billing-failure-domain/</link><guid isPermaLink="true">https://pirouter.ai/blog/billing-failure-domain/</guid><description>A Google billing account dropped from Paid Tier 3 to Tier 1 with no notice, pausing a live voice app; on NVIDIA NIM, GLM-5 retired before its replacement was callable. No status page moved, no 5xx fired. Four things to do while everything is still green.</description><pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate><category>reliability</category><category>failover</category><category>providers</category><author>Leo Kaka</author></item></channel></rss>