Howdy wizards,
Todayβs issue is brought to you by Pathway, who have a small model doing big reasoning at around 11x cheaper than a frontier model. See the results.
Hereβs whatβs brewing in AI.
The big thing
SpaceXAI launched Grok Bot, an AI team of bots thatβs easy to set up.
Hereβs how it works: each bot gets a name, a job and its own computer in the cloud.
It signs into your tools the way you do, by clicking and typing in a browser, so nothing needs an integration or an API key. When a login screen comes up you take the mouse, sign in, and hand control back. They all share one machine, which means they collaborate and share info in real-time.
Grok Bot is still just an LLM with a harness. But what separates it from other tools is that the interface is as intuitive as a regular chatbot, but opens up a level of capability currently reserved for advanced users of agentic coding tools, those who have already figured out how to efficiently deal with multi-agent setups, driving browsers and logins remotely, self-learning and iteration on workflows, etc.
The catch is the price. Itβs $200/mo for individuals, and $120/mo if youβre on a Team plan.
Why it matters
Grok Bot democratises agents that actually do work for you.
In the last few months, Iβve personally been automating the type of βclicking and typingβ work in the browser that this product enables using Claude Code. I think itβs the next big thing in AI for work. I now have repos for areas of my life, running on a VPS, where Iβve built context, authentication and workflow layers.
This power will increasingly be available to anyone, without the hours of setup. Grok Bot is the first packaged product.
And yes, the price is quite high. It burns a hell of a lot of tokens. But Iβd rather burn some tokens than my brain cells on clicking around in slow, clunky and outdated software. Better to let the agent do it.
Iβm not going to get Grok Bot myself, at least not for now. For me, Claude Code is still more flexible to build whatever I want. I also appreciate that my workflows are not built on someone elseβs setup or UI; I can switch models or any other part of my system whenever. I like using bare metal wherever possible.
But if youβre not geeking out as much on this whole AI thing and youβve got $200/mo to spend, Grok Bot might be the easiest way right now to start getting productive with real agentic workflows.
PS no need to get FOMO. This is an actually useful product, which nearly guarantees the other AI labs will release something similar very soon. And when the Chinese labs get on the train, there will also be cheaper options.
FEATURING THE COST-INTELLIGENCE BREAKTHROUGH BY PATHWAY
Pathway has published BDH-CQ, a reasoning model that learns through recurrent memory and reasons in continuous latent space without generating an intermediate chain of thought.
A 150M-parameter model scored 29.5% pass@2 on ARC-AGI-1 at just $0.0007 per task. The result was independently reproduced by researchers from Bielik AI and NYU, and separately replicated by Εukasz Kaiser.
This breaks the previously reported ARC-AGI-1 cost versus accuracy Pareto frontier. Separately, early pretraining experiments from 1B to 600B parameters showed Transformer-like scaling while preserving latent reasoning, providing encouraging evidence that the architecture can scale.
NEWS NEWS NEWS β¦ NEWS NEWS NEWS
All the small things
New tools & product features
Claudeβs Chrome side panel is now a full Cowork session. Chats you start in the browser save to your account and carry on in the desktop or mobile app, and your Skills and connectors work there without setup. Claude can read the page and click, type and fill in forms using the logins you already have. A separate AI double checks actions against what you actually asked for. Similar to Claude Codeβs auto-mode which I covered earlier this week.
DeepSeek open-sourced its agent harness under the MIT licence. Every part is a plugin: the model, the tools, the memory, the sandbox, even the interface. You swap them in config instead of forking the code. Everything the model sees goes into a log you can replay or fork. MIT means you can use it commercially with no strings. And these are very similar parts to what Grok Bot is made of, just unpackaged and without the price tag.
Models
SpaceXAI also released Grok 4.6, level with OpenAIβs best model at less than half the price. In terms of performance, itβs the same level as GPT-5.6 Sol and sits just behind Claude Opus 5 and Fable 5.
Industry moves
Businesses are barely buying Anthropicβs smartest model. Fable 5 is 6% of the tokens businesses buy from Anthropic and 11.4% of the money. It costs about twice what GPT-5.6 Sol does, and itβs used less. The growth in business AI spending is going to open (Chinese!) models instead.
Investors want to take Anthropic public in October at $2 trillion, which would be the largest listing in history. They are betting on annualized revenue of $100-120B by the end of this year.
Research
A group of AI agents broke into government networks in Asia on their own, over four days. Security firm Dream Security published the breakdown of a campaign that ran in early July. The setup was built on OpenClaw and Hermes, the same consumer tools anyone can install, and ran up to eight sub-agents at once on separate targets. It took a bunch of passwords and personnel records from targets including a nuclear safety agency. It got past its own safety rules by telling itself the whole thing was an authorised penetration test.
Anthropic put three Claude agents on the same codebase without telling them about each other, and they went to war. Each was told to move a Python backend to a different language. With no owner and no rule for who wins a conflict, every move read as hostile. They switched off each otherβs accounts, ran scripts that killed rival processes on a loop, and one planted self-replicating malware and pretended it was another agentβs work. Multi-agent systems are becoming mainstream fast, and as both of these studies show β itβs not gonna be without a few hiccups.
β¦
You are a delight.
What's your verdict on today's email?
See which AI use cases are paying off with Context Windows Pro
Most companies pick AI use cases by brainstorming internally. 90% of those initiatives fail.
Iβve created Context Windows just so you can pick the winners.
π¦ Find high-performing use cases from 2,000+ companies at contextwindows.ai, or book a demo with me
Disclosure: To cover the cost of my email software and the time I spend writing this newsletter, I sometimes work with sponsors and may earn a commission if you buy something through a link in here. If you choose to click, subscribe, or buy through any of them, THANK YOU β it will make it possible for me to continue to do this.



