Note: No affiliate links and no sponsor. I did not spend anything on any API for this, so nothing below is an invoice. Every quote comes from a vendor's own domain, opened on August 13, 2026, and every screenshot came off my own screen. Where a number is a company's own claim rather than a third-party measurement, I say so in the sentence that carries it.

Three things landed in the last two days: a frontier model launch, a 2.4-trillion-parameter open-weight release, and a quiet support-page update about watermarking. Coverage of all three moved fast. I opened the primary document behind each one, and in all three cases the document says something the coverage does not.

1. Grok 4.6 costs $2 per million tokens, until your prompt hits 200k

xAI shipped Grok 4.6 on August 12. The launch post ends with the price:

"Pricing starts at $2 per million input tokens and $6 per million output tokens. Additionally, there is a fast variant which is twice the price."

The docs site repeats those two numbers in a box with the context window.

The Grok 4.6 card on the xAI docs Models page showing Context 500k tokens, Input $2.00 per 1M tokens, Output $6.00 per 1M tokens, with no mention of any pricing threshold docs.x.ai/docs/models on August 13. Context and price in the same box: 500k tokens, $2.00 in, $6.00 out.

Read those together and you get half a million tokens of frontier context at two dollars a million. The pricing page on the same site splits the row in two.

The xAI pricing table row for grok-4.6 showing 500k context, a note reading Long context greater than or equal to 200k tokens, short context prices of $2.00 input, $0.50 cached and $6.00 output, and long context prices of $4.00 input, $1.00 cached and $12.00 output Same docs site, Pricing page. Short context and Long context are separate columns, and the threshold sits under the model name.

Past 200k prompt tokens the rate becomes $4.00 in and $12.00 out. Cached input goes from $0.50 to $1.00. And then there is one sentence, in the smallest grey type on the page, that decides what you actually pay:

"Models with long context pricing bill the long context rates for all tokens in a request once its prompt reaches the model's long context threshold."

The higher rate is not charged on the tokens above 200k. It is charged on every token in the request, once the prompt crosses the line.

A 250,000-token prompt with 8,000 tokens of output, priced the way the launch post reads: 250,000 × $2 / 1M = $0.50, plus 8,000 × $6 / 1M = $0.048, so $0.55. The same request as the table bills it: 250,000 × $4 / 1M = $1.00, plus 8,000 × $12 / 1M = $0.096, so $1.10.

The step is sharper than the doubling suggests. A request with 199,999 input tokens costs $0.40. A request with 200,001 costs $0.80. One token, twice the bill. Filling the advertised 500k window once costs $2.00 in input, not the $1.00 that "500k" and "$2 per million" imply when you read them in the same box.

Step chart of what one Grok 4.6 request costs by prompt size. A shallow blue line rises to forty cents just under two hundred thousand input tokens, an orange dashed line jumps vertically to eighty cents at the two hundred thousand token threshold, and a steeper blue line continues to two dollars at five hundred thousand tokens Input cost of a single request, computed from the published rates. The jump at 200k is the whole story.

Two things I am not claiming. This tiering is not new and not unique to 4.6: the same page shows a 200k threshold on grok-4.5, grok-4.3, grok-build-0.1, and the grok-4.20 family. And the launch post's "fast variant which is twice the price" has no matching row for 4.6 in the table I captured, so I am flagging that gap rather than filling it.

On where the model lands, xAI says it "matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index." Artificial Analysis, which is the third party doing that measuring, puts a number on it: Grok 4.6 scores 61, tied with GPT-5.6 Sol, behind Claude Opus 5 at 63 and Claude Fable 5 at 62. The same writeup measured $0.84 per task and reports Grok 4.6 finishing tasks in about 53 turns and 0.5B input tokens, against about 103 turns and 2.0B for Claude Opus 5. Efficiency is doing the work in that per-task number, which is a different claim from a low per-token rate.

2. Qwen3.8's license does not ban the US, the EU, the UK, or Korea. A different model's license does.

Alibaba published open weights for Qwen3.8-2.4T-A95B on Hugging Face this week, and a claim spread with it: that the license forbids use in the United States, the European Union, the United Kingdom, and South Korea. If you build on open weights, that sentence decides whether you can touch the model at all.

Here is what each of the two licenses actually says, side by side.

Comparison table titled which license restricts which region. Qwen3.8-2.4T-A95B has no regional restriction and no territory clause, restricting only naming above one hundred million monthly active users or twenty million dollars monthly revenue and requiring a separate license for Model as a Service businesses above fifty million dollars. MiniMax H3 excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America The four-region list exists. It is in the other file.

I opened the Qwen LICENSE file. It is sixteen lines.

The full LICENSE file for Qwen3.8-2.4T-A95B on Hugging Face, 3.39 kB and sixteen lines, showing an MIT-style permission grant, a naming requirement clause, a Model as a Service clause, and a warranty disclaimer, with no geographic restriction anywhere The entire license, 3.39 kB. There is no territory clause in it.

What the file actually restricts is commercial scale, in two places. Products above 100,000,000 monthly active users or US$20,000,000 monthly revenue have to display the model name prominently. Businesses running "Model as a Service or AI Work Assistant" with aggregate revenue above US$50,000,000 over twelve consecutive months need a separate license from Qwen first. Beyond that the grant reads like an MIT license: use, copy, modify, merge, publish, distribute, sublicense, sell, deploy, host, fine-tune, and create derivative works.

The four-region list is real, and it belongs to a different model. MiniMax's H3 license, released August 2, defines its terms this way:

Line 10 of the MiniMax H3 Community License Agreement on Hugging Face reading: Excluded Territories means the European Union, the United Kingdom, the Republic of Korea and the United States of America MiniMax H3's LICENSE, line 10. This is the clause people have been attributing to Qwen.

That agreement grants rights "solely within the Applicable Territory," defines Applicable Territory as worldwide excluding the Excluded Territories, and adds that you may not use or display the works or their outputs outside it. So a restriction that exists on one Chinese open-weight release got transferred, in the retelling, onto a different one.

While the license was being described wrong, the model was too. The card is explicit about what these weights are:

"Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc."

So this is not Max going open source, it is the base model Max is built on. And:

"Qwen3.8-2.4T-A95B is a text-only model that requires thinking mode for all interactions. Multimodal inputs are not supported, and thinking cannot be disabled."

The multimodality in the coverage belongs to the API product, not to the weights you can download. The weights are 2.4T total with 95B activated, 262,144 tokens of native context, extensible to 1,010,000.

3. Claude now watermarks its output, and Anthropic says the watermark is not proof

Anthropic did not announce this in a blog post. It appeared as a support article, updated this week, which is itself worth noticing given how the story travelled.

The mechanism, in Anthropic's words, is that Claude "weaves an imperceptible watermark directly into the text itself," and that mark "will travel with the text when it's copied and pasted elsewhere, and may persist through some editing." Files get C2PA signed provenance metadata instead. Models launched on or after August 2, 2026 support marking at launch, older models are being brought in, and the coverage spans the API, Claude, Claude Code, Claude Cowork, and Claude Tag. Marking applies "wherever Claude is offered, worldwide," which means a rule written for the EU AI Act is being applied globally.

The part that decides how much this can be used against a student or an employee is on the same page:

"A detected mark provides a signal that content was processed by Claude, but is not fully conclusive."

And in the other direction, content generated by Claude may carry no detectable mark if the text was heavily edited, paraphrased, translated, mixed into other writing, or is simply very short, and metadata can be stripped from files. A mark can also appear on text whose ideas and words came from somewhere else, because Claude only had to touch it.

So the honest reading is: Anthropic can mark, Anthropic says the mark is a signal rather than proof, and absence of a mark means nothing at all. The support page points to forthcoming technical documentation for detection, and I could not find a public detection tool offered anywhere on it as of August 13, which means that for now nobody outside Anthropic can check a given piece of text. No independent evaluation of how well the watermark survives editing exists yet either, two days in. If you see a confident claim that a paragraph "was written by Claude," ask what tool produced that verdict.

What the three have in common

In each case the vendor's own document is more precise than the story built on it, and in each case the precision is the part you need: the threshold that doubles your bill, the clause that is not in the license, the sentence saying the evidence is not conclusive. All three documents are public, unauthenticated, and took a few minutes each to open.

Sources, all read on August 13, 2026: xAI's Grok 4.6 launch post and the models and pricing pages on docs.x.ai; Artificial Analysis's Grok 4.6 writeup; the Hugging Face model card and LICENSE for Qwen/Qwen3.8-2.4T-A95B; the LICENSE for MiniMaxAI/MiniMax-H3; and Anthropic's support article on how Claude marks AI-generated content.