Blog

Kimi K3 license, commercial use: the $20M MaaS clause, the 100M-MAU badge, and the 30% Moonshot wants from Azure, AWS and Google

K3's weights are free; serving starts at 8× GB300, so a host stays. Reuters: Moonshot wants up to 30% from Azure, AWS and Google — token auditing is unresolved.

Linden Kern11 min read
Four-stage diagram for Kimi K3: 2.8T weights free to download → host needs 8× GB300 or more → reported revenue share of up to 30% → token auditing unresolved. Below it, the three license thresholds: $20M over 12 months, 100M MAU or $20M a month, internal use and certified partners exempt

Three sentences first. The Kimi K3 License is MIT-shaped with two commercial thresholds on top — a $20 million revenue line that sends a Model-as-a-Service business to a separate agreement (§2), and a 100-million-user or $20-million-a-month line that puts “Kimi K3” in your UI (§3) — plus exemptions for internal use and Moonshot’s own channels (§4); the table below has the cases. The weights are free to download, but the official vLLM recipe starts at eight GB300s and says multi-node for production — a host is still in the path. And according to Reuters on 26 August, citing three people familiar with the talks, Moonshot is asking Microsoft, Amazon and Google for up to 30% of the revenue their clouds generate from K3 — early-stage talks that may not conclude, with the revenue split, data access and “auditing token usage” all unresolved.

That last item is what this piece is about. On 24 August we wrote that the layer between developers and models had been priced twice in one month; this is the third pricing, and it points the other way.

What the Kimi K3 license actually says: the commercial conditions, in one table

The LICENSE file is 3,065 bytes. The grant is broad — use, copy, modify, distribute, sublicense, sell, deploy, fine-tune, all named. §1 applies to everyone and never switches off: the copyright and permission notice “shall be included in all copies or substantial portions of the Software,” plus compliance with law. §4 exempts internal use and Moonshot’s own channels from §2 and §3 only — it says nothing about §1. The table is therefore about commercial conditions; attribution rides along with any copy you pass on, fine-tuned checkpoints included. Our reading, not legal advice.

Your use of K3§2 separate agreement§3 “Kimi K3” on the UIVerdict (commercial conditions)
Internal tools; outputs never reach third partiesNot applicableNot applicableNone (§4 (a)); §1 still applies to copies you distribute
Embedded in your product, model capability confined to specific featuresOutside the MaaS definitionAbove 100M MAU or $20M/monthBadge only, at scale
Relaying requests to a model someone else hosts (a router or gateway)Outside the MaaS definition (§2 (b))SameBadge question only
Fine-tuning and redistributing the weights, without selling inferenceOutside the MaaS definitionSame§1 travels with the checkpoint
Hosting the weights and selling inference or fine-tuning to third partiesRequired once aggregate revenue with affiliates exceeds $20M over any 12 monthsSameSeparate agreement above the threshold
Access via Moonshot’s official products or certified inference partnersExemptExempt§4 (b)

§2, verbatim:

“Model as a Service” means giving a third party access to language model inference or fine-tuning (e.g., via API) in a manner that allows such third party to exercise meaningful control over the inputs, parameters, or training data. This does not include (a) end-user products with model capabilities solely embedded within specific features or harnesses, or (b) mere relaying of requests to models hosted by others.

If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

§3 and §4:

If the Software (or any derivative works thereof) is used for any of the Licensee’s commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, “Kimi K3” must be prominently displayed on the user interface of such product or service.

The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; or (b) any use of the Software accessed through Moonshot AI’s official products or certified inference partners.

The $20 million in §2 is aggregate revenue of the licensee and its affiliates, not revenue from K3: a large company that puts K3 behind an API crosses the line on day one however little K3 earns. Crossing it obliges you to sign a separate agreement; no rate appears in the text. The two $20 million figures are also different quantities — §2 a twelve-month total, §3 a monthly one.

§2 (b) carves out “mere relaying of requests to models hosted by others,” so a router or gateway that does not run the weights falls outside the definition.

“Certified inference partners” in §4 (b) has no definition and no list — TrendingTopics noted the same when the weights dropped. An exemption whose beneficiaries the license does not name is what connects it to the Reuters story.

Downloadable is not runnable: 2.8T parameters, at least 8× GB300

From the model card: 2.8 trillion total parameters, 104 billion active, 16 of 896 experts selected per token, 93 layers, a 1,048,576-token context, weights shipped in MXFP4 from quantization-aware training.

The argument “it is open-weight, so we can run it ourselves and skip both the license and the cloud” stops at the serving recipes. The vLLM page’s prerequisites:

Hardware: At least 8x GB300. Multi-node for real production traffic.

The ROCm line says at least 8× MI355X or MI350X, and the MXFP4 footprint on the same page is 1,680 GB, footnoted as an estimate until the checkpoint is published. The SGLang cookbook spells out node layouts per accelerator, and the card count moves with how new the silicon is: B300 and MI355X at 1 node × 8 cards, GB300 at 2 × 4 — eight cards either way — but H200 and B200 at 2 × 8, and H100 at 4 × 8, which is thirty-two. Every cell in that panel is marked Final Verification In Progress: the recipe runs, but verification on final weights and current code is still open. Eight cards is the floor only if they are this year’s; a fleet of H100s needs four times that.

Node × GPU layouts for serving Kimi K3 in the SGLang and vLLM official recipes
SGLang cookbook layouts (H100 4×8 / H200 2×8 / B200 2×8 / B300 1×8 / GB300 2×4 / MI355X 1×8) and the vLLM recipe floor of 8× GB300, read on 2026-08-29.Source: SGLang cookbook — Kimi-K3, vLLM Recipes — moonshotai/Kimi-K3

Reuters makes the same point: analysts say few customers are likely to run a 2.8-trillion-parameter system on their own infrastructure. Supply is not slack either — Moonshot paused new subscriptions after the July launch because its own GPUs could not keep up (GIGAZINE, RuntimeWire).

Open weights and no host are two different propositions. K3 has the first, not the second.

The 30% ask runs the platform split backwards

Reuters’ sources are “three people familiar with the talks”; the figure is “up to a 30% share of revenue generated from K3-related services”; the status is “at an early stage and there is no certainty they will result in agreements.” Moonshot did not respond; Microsoft, Google and AWS declined to comment. Nothing says a deal exists, or that K3 is coming to any of the three catalogs.

The direction is what deserves attention. The Next Web: “It also inverts the usual arrangement, in which the platform takes the cut, and the developer receives the remainder.” Here the party giving away the weights asks the platform for a percentage. RuntimeWire calls the 30% “a ceiling rather than an agreed rate”; Reuters says it matches terms Moonshot has outlined to large customers.

Reuters reports similar agreements already signed with smaller cloud platforms, a July disclosure from Chinasoft International with the split unpublished, and one more line at the end: Alibaba is seeking revenue-sharing agreements with major users of its own new open-source model. Weights free, revenue collected where the weights run — the shape has more than one data point.

Now connect this to §4 (b). If the three clouds become “certified inference partners,” their customers fall outside §2 and §3 entirely. The license threshold faces the customer; the revenue share faces the host. In our reading, two collection windows for one economic claim — an inference, labelled as one.

The politics get one paragraph. Reuters reports that Treasury Secretary Scott Bessent said in July he might add Moonshot to a trade blacklist, that U.S. officials have accused the company of distilling from Anthropic’s Fable model and of illegally acquiring Nvidia chips, and that Moonshot rejects the distillation claim, attributing K3’s gains to original architectural changes. None of it is settled, and these talks may end on policy rather than terms. This piece stops there.

Cartoon: a "FREE — TAKE ONE" stall with one WEIGHTS packet, beside a house-sized machine plated 8 × GB300
The weights are free. The place that runs them is not.

Who counts the tokens? Three ways to audit usage between a host and a model owner

Back to the third unresolved item. Reuters added a gloss: “Tokens are units of text processed by AI models, and measuring their use is central to calculating revenue under usage-based billing.” TNW went further: the method has to work “in an arrangement where the party doing the counting is the party paying the share.”

Between a host that meters and an owner paid a percentage, an audit takes one of three shapes, each with a cost.

ApproachWhat happensWhere the cost lands
A. Host-side metering plus sampled auditHost keeps per-request records (input, output and reasoning tokens, model ID, timestamp); owner pulls a window and reconcilesStorage on the host, audit effort on the owner, trust in the host’s ledger
B. Independent metering on both sides, reconciled periodicallyGateway count and model-side count (tokenizer or engine usage) captured separately; the monthly difference has to be explainedTwo meters to run, and a recurring investigation into why they differ
C. Verifiable meteringRecords carry a hash chain or append-only ledger the owner can read but not alterHighest build cost, and a head-on collision with what the owner may see

Whichever shape you pick, one thing has to be settled first: what is one token. K3 makes this sharp — the model card says it “always has thinking enabled, and will return reasoning_content.” Four definitions have to be written down before any meter means anything:

  • Do reasoning tokens count as output for the split?
  • Whose conversion table applies to image input?
  • Do input tokens served from a prefix cache count as “processed”?
  • Under preserved thinking history — the full prior assistant message re-sent every turn inside a 1M-token window — is re-sent input counted each time?

These are contract definitions, not engineering ones. If they drift, approach B produces a monthly difference nobody can explain.

Our control plane is designed to meter every request independently of the upstream and reconcile against its bill — designed, not a shipped feature. Where the count is taken is why seven endpoints, one name, two prices and Shared team wallet for LLM APIs went the same way: an invoice is only as good as the place it was measured.

My reading is that when one of Reuters’ sources named “auditing token usage” as unresolved, these definitions are part of what remains open. The 30% is the easy number; the denominator it multiplies is the hard one.

Where this leaves the distribution layer

The first two pricings of this layer — the hub you fetch weights from, the router you send requests through, both put on the market within a month — were about who owns it. This one is about who takes a cut of what flows across it. What is being valued is the same in all three: not the model, but the path to it.

Three things are worth watching, none of which requires predicting whether a deal closes.

  1. Whether a list of certified inference partners is ever published. Until it is, any procurement leaning on the §4 (b) exemption is leaning on air. If you host, apply §2 to your aggregate revenue; if you relay, confirm you are inside §2 (b).
  2. How the data-access clause is written. It is the second of Reuters’ unresolved items, and it decides which of the three audit shapes is even possible. Without B or C, “audit” means reading the host’s books.
  3. Who defines the counting rules. Reasoning tokens, image tokens, cache hits, re-sent context. If these are not in the contract, the denominator moves every month.

The license takes ten minutes to read. What is left afterwards is not a licensing question but a metering one. Until someone settles who counts the tokens and who can check the count, “open weights, so you are free” is true for anyone who owns eight GB300s.