Howdy wizards,

Here’s what’s brewing in AI.

The big thing

Moonshot’s Kimi K3 is a Chinese open model that just caught up with OpenAI and Anthropic.

How good is this model? On the most reputable independent index (Artificial Analysis), it sits a hair behind the two best models in the world: Anthropic’s Fable 5 and OpenAI’s GPT-5.6.

And the cost story is huge. Kimi K3 offers frontier performance at a fraction of what Fable 5 and GPT-5.6 cost ($3/million input tokens and $15/million output).

It’s probably a big model: 2.8 trillion parameters. I say probably because big is relative, and we don’t really know how big the leading American models are, because they’ve stopped announcing their size.

Also, it’s an open weights model, which means you can run your own copy instead of renting it from the company. That does not mean β€œruns on your laptop” though. Far from it, actually: according to Moonshot, it takes 64 high-end GPUs to do this bad boy full justice.

Why it matters

So you thought Chinese models trailed the US labs by six to twelve months? That’s what Dario (Amodei, not me) said a while ago. But Kimi K3 came out less than two months after Fable, so the gap seems to have shrunk.

A top-three model is now open and cheap. Models are turning into commodities. And it’s putting pressure on American labs.

For all you Claude users, here’s the good news: the pressure from this cheap, Chinese powerhouse of a model, plus OpenAI’s new GPT-5.6 being included in the regular ChatGPT plan, is forcing Anthropic to keep its best model cheap. They’ve been saying Fable 5 will become pay-as-you-go very soon, but they’ve now decided to make it included in their Max and Enterprise plans.

One more piece of good news, this time for anyone whose calendar is full of meetings. Which is where Granola, this week's sponsor, comes in:

IN PARTNERSHIP WITH GRANOLA

Your meetings are your richest source of context. Use Granola for a while and you’re sitting on a goldmine: months of decisions, discussions, and commitments, all captured.

But how often do you actually go back and search through them? Basically never. It’s a hassle, and your notes have no way of talking back to you.

Granola Chat fixes that: an AI chat that works across all your notes, personal and team. Ask it a question and it searches everything: your notes, your Team Space, and privately shared notes.

Quick question? Fast answer. Deeper analysis, like spotting trends across fifty sales calls? It handles that too, and every answer links back to the source with inline citations.

All that context you’ve been building compounds now. Chat turns your archive into an always-on analyst that already knows your business.

NEWS NEWS NEWS ❦ NEWS NEWS NEWS

All the small things

Industry moves

  • Anthropic is now worth more than OpenAI. It just closed a $65 billion round at a $965 billion valuation, edging past OpenAI’s $852 billion, and bankers are lining up a possible IPO later this year. DeepSeek is raising in China ahead of its own listing too. AI, bubble or not, is about to hit the public markets.

  • Over 200 economists signed a warning that AI will affect jobs faster than we are ready for. 16 Nobel laureates signed, plus researchers from Google, OpenAI, and Anthropic. The claim: past shifts gave societies decades to adjust, AI might give us a few years. It’s heavy on names, but light on what to actually do. But the names are the point (at this point).

  • Demis Hassabis wants a referee for frontier AI, and he wants it running this year. The DeepMind CEO proposed a US-led body, funded by the industry, that would test new models for bio, cyber, and deception risks 30 days before release. Voluntary at first, then required to reach the US market. The obvious catch: having a watchdog paid for by the labs it is meant to watch.

  • New York became the first state to pause new data centers. The freeze covers anything over 50 megawatts, for up to a year, while the state writes power and water rules. Elon Musk quietly bought a gas-turbine company to make his own electricity for Grok. And SK Hynix warned the memory-chip shortage could run through 2030. The bottleneck is moving from software to power plants and factories.

  • Rumour mill: OpenAI’s first gadget is reportedly a screenless speaker that acts like a ChatGPT you can talk to at home. Bloomberg says the Jony Ive-designed device is small, rechargeable, and screen-free, with cameras, sensors, and a bit of movement, aimed at 2027. It studies things like your email to get to know you. An always-listening camera-speaker is quite the privacy trade, and it happens to be the same device Apple is suing to block.

Models

  • Mira Murati’s lab shipped its first model, Inkling, and made it open. The $12 billion Thinking Machines Lab released a multimodal open-weight model with a dial that trades thinking time for cost. It trails the top Chinese open models, and was reportedly trained partly on outputs from an older Kimi model.

  • Google delayed its flagship Gemini 3.5 Pro over weak coding performance. The model is reportedly months behind schedule, and the delay has Google’s own engineers worried about falling behind. When open models are speeding up and your flagship is slipping, the market notices; Google stock dipped about 4% on the news.

New tools & product features

  • OpenAI made a $230 keypad for bossing around its coding agent. The Codex Micro is a little mechanical control pad: keys that light up when the agent needs a decision, a joystick to switch jobs, a dial for how hard it should think. Actually a really cool looking toy. I’m tempted. Looking at the white box packaging though β€” 100% Apple vibes a week after they got sued by them πŸ˜†

Research

  • OpenAI built an AI that attacks its own models to make them safer. GPT-Red writes its own adversarial prompts to hunt for holes; OpenAI says training against them cut GPT-5.6’s prompt-injection failures roughly sixfold. Prompt injection (hidden text hijacking an AI agent) is still one of the field’s ugliest unsolved problems. Whether GPT-Red catches the attacks nobody has thought of yet is an open question.

  • Anthropic found that Claude’s personality shifts with the language you use. Anthropic admits it does not know why. Across 309,000 chats and 20 languages: Dutch drew more admissions of error, Hindi more warmth, English more detail. So when you need AI to admit it messed everything up, just switch to Dutch…

❦

You are a delight.

See which AI use cases are paying off with Context Windows Pro

Most companies pick AI use cases by brainstorming internally. 90% of those initiatives fail.

I’ve created Context Windows just so you can pick the winners.

🟦 Find high-performing use cases from 2,000+ companies at contextwindows.ai, or book a demo with me

Disclosure: To cover the cost of my email software and the time I spend writing this newsletter, I sometimes work with sponsors and may earn a commission if you buy something through a link in here. If you choose to click, subscribe, or buy through any of them, THANK YOU – it will make it possible for me to continue to do this.

Keep Reading