Howdy wizards,

Today’s issue is brought to you by Pathway, who have a small model doing big reasoning at around 11x cheaper than a frontier model. See the results.

Here’s what’s brewing in AI.

The big thing

SpaceXAI launched Grok Bot, an AI team of bots that’s easy to set up.

Here’s how it works: each bot gets a name, a job and its own computer in the cloud.

It signs into your tools the way you do, by clicking and typing in a browser, so nothing needs an integration or an API key. When a login screen comes up you take the mouse, sign in, and hand control back. They all share one machine, which means they collaborate and share info in real-time.

Grok Bot is still just an LLM with a harness. But what separates it from other tools is that the interface is as intuitive as a regular chatbot, but opens up a level of capability currently reserved for advanced users of agentic coding tools, those who have already figured out how to efficiently deal with multi-agent setups, driving browsers and logins remotely, self-learning and iteration on workflows, etc.

The catch is the price. It’s $200/mo for individuals, and $120/mo if you’re on a Team plan.

Why it matters

Grok Bot democratises agents that actually do work for you.

In the last few months, I’ve personally been automating the type of β€œclicking and typing” work in the browser that this product enables using Claude Code. I think it’s the next big thing in AI for work. I now have repos for areas of my life, running on a VPS, where I’ve built context, authentication and workflow layers.

This power will increasingly be available to anyone, without the hours of setup. Grok Bot is the first packaged product.

And yes, the price is quite high. It burns a hell of a lot of tokens. But I’d rather burn some tokens than my brain cells on clicking around in slow, clunky and outdated software. Better to let the agent do it.

I’m not going to get Grok Bot myself, at least not for now. For me, Claude Code is still more flexible to build whatever I want. I also appreciate that my workflows are not built on someone else’s setup or UI; I can switch models or any other part of my system whenever. I like using bare metal wherever possible.

But if you’re not geeking out as much on this whole AI thing and you’ve got $200/mo to spend, Grok Bot might be the easiest way right now to start getting productive with real agentic workflows.

PS no need to get FOMO. This is an actually useful product, which nearly guarantees the other AI labs will release something similar very soon. And when the Chinese labs get on the train, there will also be cheaper options.

FEATURING THE COST-INTELLIGENCE BREAKTHROUGH BY PATHWAY

Pathway has published BDH-CQ, a reasoning model that learns through recurrent memory and reasons in continuous latent space without generating an intermediate chain of thought.

A 150M-parameter model scored 29.5% pass@2 on ARC-AGI-1 at just $0.0007 per task. The result was independently reproduced by researchers from Bielik AI and NYU, and separately replicated by Łukasz Kaiser.

This breaks the previously reported ARC-AGI-1 cost versus accuracy Pareto frontier. Separately, early pretraining experiments from 1B to 600B parameters showed Transformer-like scaling while preserving latent reasoning, providing encouraging evidence that the architecture can scale.

NEWS NEWS NEWS ❦ NEWS NEWS NEWS

All the small things

New tools & product features

  • Claude’s Chrome side panel is now a full Cowork session. Chats you start in the browser save to your account and carry on in the desktop or mobile app, and your Skills and connectors work there without setup. Claude can read the page and click, type and fill in forms using the logins you already have. A separate AI double checks actions against what you actually asked for. Similar to Claude Code’s auto-mode which I covered earlier this week.

  • DeepSeek open-sourced its agent harness under the MIT licence. Every part is a plugin: the model, the tools, the memory, the sandbox, even the interface. You swap them in config instead of forking the code. Everything the model sees goes into a log you can replay or fork. MIT means you can use it commercially with no strings. And these are very similar parts to what Grok Bot is made of, just unpackaged and without the price tag.

Models

  • SpaceXAI also released Grok 4.6, level with OpenAI’s best model at less than half the price. In terms of performance, it’s the same level as GPT-5.6 Sol and sits just behind Claude Opus 5 and Fable 5.

Industry moves

  • Businesses are barely buying Anthropic’s smartest model. Fable 5 is 6% of the tokens businesses buy from Anthropic and 11.4% of the money. It costs about twice what GPT-5.6 Sol does, and it’s used less. The growth in business AI spending is going to open (Chinese!) models instead.

  • Investors want to take Anthropic public in October at $2 trillion, which would be the largest listing in history. They are betting on annualized revenue of $100-120B by the end of this year.

Research

  • A group of AI agents broke into government networks in Asia on their own, over four days. Security firm Dream Security published the breakdown of a campaign that ran in early July. The setup was built on OpenClaw and Hermes, the same consumer tools anyone can install, and ran up to eight sub-agents at once on separate targets. It took a bunch of passwords and personnel records from targets including a nuclear safety agency. It got past its own safety rules by telling itself the whole thing was an authorised penetration test.

  • Anthropic put three Claude agents on the same codebase without telling them about each other, and they went to war. Each was told to move a Python backend to a different language. With no owner and no rule for who wins a conflict, every move read as hostile. They switched off each other’s accounts, ran scripts that killed rival processes on a loop, and one planted self-replicating malware and pretended it was another agent’s work. Multi-agent systems are becoming mainstream fast, and as both of these studies show β€” it’s not gonna be without a few hiccups.

❦

You are a delight.

See which AI use cases are paying off with Context Windows Pro

Most companies pick AI use cases by brainstorming internally. 90% of those initiatives fail.

I’ve created Context Windows just so you can pick the winners.

🟦 Find high-performing use cases from 2,000+ companies at contextwindows.ai, or book a demo with me

Disclosure: To cover the cost of my email software and the time I spend writing this newsletter, I sometimes work with sponsors and may earn a commission if you buy something through a link in here. If you choose to click, subscribe, or buy through any of them, THANK YOU – it will make it possible for me to continue to do this.