Wizard sitting in the open back of a van beside a desert road, arms crossed, a steaming mug and a laptop beside him

Howdy,

and welcome to the 484 new wizards who joined this month.

A hotel let me store my suitcase for 2 weeks in their mostly unmanned reception. I made a judgement call: 95% chance it’s still there when I come back. Worldward’s AI agent said 78%. And holy luggage, was it there. So we were both correct (I was just a little more confident). Want to record your decisions in an AI journal? Try Worldward.

Here’s what’s brewing in AI.

The big thing

What do the new Claude Opus 5.5 and GPT-6 Sol have in common? Shorter answers, fewer lies, cheaper prices.

It’s been days since the “slow down” outrage from Anthropic’s Dario Amodei, OpenAI’s Sam Altman and other AI leaders.

You’d think they’d lay low for a few weeks? Nah.

The two labs everyone’s watching are already at it again.

This time with an update of their two workhorse models.

Let’s get into it:

Opus 5.5 and GPT-6 Sol aren’t their top-top models; that’s Fable for Anthropic, Astra for OpenAI.

The Opus and Sol series are the models they want you to default to for most things. In Norway we say: å skyte spurv med kanoner (to shoot sparrows with cannons). That’s what the labs want to avoid people doing, so they make these mid-tier models with mid-level effort levels the default, so you don’t blow all their precious tokens on writing your damn emails. They’re designed to be the right sized guns for daily work.

“Ok, cool, new models again. So what’s the difference between these and whatever was before them?”

  1. They make less sh*t up. Early testers of Opus 5.5 report it’s way more accurate than its predecessor. OpenAI says the new Sol makes half as many factual errors.

  2. They get to the point faster. I’m talking about the length of the answers they give. Opus 5.5 is 40% less wordy, while OpenAI says Sol has fewer low-value details, less jargon and weird ways to say things.

  3. They’re 20-50% cheaper than the previous versions. Relative to each other: Opus 5.5 is double the price of Sol. And for perspective, they’re still about 10-15x pricier than DeepSeek’s comparable model.

Judging by developer sentiment, I’d say people are most excited about the new Opus model. METR (an independent evaluation lab) also calls Opus 5.5 an incremental improvement over Fable. So it’s even better than what’s supposed to be a higher-tier model.

PS OpenAI also launched GPT-6 Luna, the fastest & cheapest version in the series. Relevant if you’re building high-volume use cases.

Why it matters

Better models at lower prices and you don’t even have to jump between companies* for this release since they both came at the same time. I just started using Opus 5.5 so it’s early to say what I think, but so far it kinda feels like this 🕺🏻

The bigger thing in all of this is that the models’ prices keep getting lower while capability climbs.

How much is the actual cost curve, and how much is these companies being forced to push prices by the competition from open weight models? Open is getting really popular as more devs realise it’s much cheaper for almost the same thing.

  • Both companies are feeling the pressure to have great models at an affordable price point. They know that deep down, companies actually don’t care what the model has to say about Tiananmen Square because they’re using AI to churn out code. They are in it for profit.

  • That’s also why every time these new models arrive they’re at a discount. The more natural thing would be that a new product is more expensive than the last, wouldn’t it?

  • Effort levels are a smart tool the big labs already have in place so that instead of running off asking DeepSeek for answers to 90% of your questions, you can e.g. use Opus 5.5 with low or medium effort.

*side note, if you FOMO into switching back & forth between these providers. I’ve followed this space for 3 years now. Not once have they not been one-upping each other on a weekly basis. If you just wait a bit and they’ll have parity soon enough.

--

The new models make fewer things up.

But they have no idea what was said in your meetings.

Thank god Granola does:

IN PARTNERSHIP WITH GRANOLA

Granola Chat: prompt bars reading What did I promise in my last call with Rob, Give me background of this person before my 2pm, and What did we decide about pricing last month

100? 500? More?

Your brain can’t hold all that context and surface the right information at exactly the right moment. That’s why Granola Chat exists.

Let’s say you need a quick refresher on where a project is at, mid-meeting. Use a Granola recipe like /Backstory. It searches across your notes and gives you a quick recap at lightning speed.

Need a deeper dive? Chats with Granola combine insights from across meetings, notes, attachments and the web (with links back to every source, of course).

Whether you’re writing your own prompts or choosing from a library of pre-made Recipes, Granola Chat is your PA, chief of staff, work bestie and brain trust, all rolled into one.

NEWS NEWS NEWS ❦ NEWS NEWS NEWS

All the small things

Industry moves

  • Google is launching its first AI chips into orbit on October 1. A fridge-sized satellite with four of Google’s TPU chips goes up to see whether they survive launch, radiation and heat. Solar panels get up to 8x the power in orbit, and data centres down here are running into power limits. it's a bird it's a plane it’s Google’s orbital datacenter

  • Meta’s Muse: blocked from shopping on Amazon, but invited to shop at Shopify. Different plays on how to handle people using agents to shop for them from the big ecom platforms. Amazon say they don’t like that Muse browses without saying it’s an agent. Amazon’s ad business needs you browsing. An agent that just buys skips it.

  • OpenAI launched Sponsored Agents: ads in ChatGPT you can talk to. Click one and you get a separate, labelled chat with an agent the advertiser built. OpenAI’s demo is a dining table: ask how many it seats, then follow the link to buy. OpenAI is testing it with selected US advertisers. Trending: affiliate marketing on agentic steroids.

Models

  • An ex-OpenAI researcher’s startup released Jev, a model that makes decisions instead of writing text. Give it text and the possible answers, and it picks one in under half a second. It only answers from your list, so it’s for the small judgment calls inside software: spam or not, which folder, which model gets the job. Hacker News went crazy over this, more upvoted than both Opus 5.5 and GPT-6.

New tools & product features

  • Claude Cowork is now part of the regular chat, and keeps working after you close your laptop. You don’t have to decide where a task belongs any longer; what Cowork can do is available from the regular chat.

Cold brew

In this week’s Can You Spot the Vibe Coder?

that’s actually ME last week, attempting not to interrupt the 12 agents I spawned right before my flight.

❦

You are a delight.

See which AI use cases are paying off with Context Windows Pro

Most companies pick AI use cases by brainstorming internally. 90% of those initiatives fail.

I’ve created Context Windows just so you can pick the winners.

🟦 Find high-performing use cases from 2,000+ companies at contextwindows.ai, or book a demo with me

Disclosure: To cover the cost of my email software and the time I spend writing this newsletter, I sometimes work with sponsors and may earn a commission if you buy something through a link in here. If you choose to click, subscribe, or buy through any of them, THANK YOU – it will make it possible for me to continue to do this.