undefined nº 02
What is happening with Anthropic?
Hey,
We closed issue one wondering how many new models this one would have to cover. Seven, as it turns out, and somehow none of them was the story of the fortnight.
// anthropic had a rough two weeks
Anthropic announced a permanent 25% increase to Claude Code’s weekly limits, starting September 14. Users did the math and found the other number: the increase replaces a temporary 50% boost that has been running since May, so measured against what people have today it is a 17% cut.
Both numbers are true. The difference is which one you put in the headline.
Anthropic deleted the original thread and reposted it, and the second version says it in plain words:
Compared to today, this works out to a 17% reduction in weekly limits on Claude Code.
It landed on dry ground. The “20x” in the $200 Max plan had already been dissected at length, because the multiplier applies inside five-hour windows rather than to the weekly ceiling, and a class action filed in June alleges that Max 20x subscribers get something nearer 6-8x the Pro plan than the 20x on the label. Then Shopify’s CEO Tobi Lütke publicly floated banning Claude Code until it reads AGENTS.md like every other agent, which for a team running several coding agents across one large codebase is a maintenance tax rather than a matter of taste.
None of this is about model quality. Claude is still very good, and credit where it is due, Anthropic’s own people engage with the criticism honestly and in public. The problem is simpler: customers should be able to calculate what they bought, and right now a lot of Claude subscribers can’t.
OpenAI noticed. Codex spent August resetting limits and shipping harness fixes it reckons are worth 10-50% more work per quota dollar, which is the opposite direction.
// openai breaks up with cursor
Different company, same lesson. After SpaceX closed its $60 billion acquisition of Anysphere, Cursor’s parent, OpenAI announced it will wind down Cursor’s direct access to its models, with a proposed shutoff on November 12. The stated reason is that OpenAI cannot be confident SpaceX will respect its terms, given its previous experience with Musk’s companies. Developers will still be able to bring their own API keys.
Cursor’s answer was basically a shrug: CEO Michael Truell says OpenAI models account for about 5% of its traffic.
That number is the story. Composer, the model Cursor built for itself, sits on top of Moonshot’s open-weights Kimi K2.5, and the rest of the stack is spread across several providers, so losing a closed lab is an inconvenience rather than an extinction event. Anthropic, for its part, was out within hours saying publicly that Cursor has been a trusted partner since Sonnet 3.5.
Two stories, one conclusion: when your stack depends on a single closed provider, their decisions become your roadmap. Open weights are how Cursor earned its shrug.
Cursor’s response, and the 5% figure
// quick hits: the flash wars
Seven models landed in two weeks, and the interesting part isn’t the number but where the competition has moved, which is fast, cheap inference. Four of the seven are Flash-class.
- DeepSeek v4 Flash Vision Exp. Vision on the Flash backbone, with the weights on Hugging Face under MIT since August 31. Same 1M context, and every image bills at up to 384 tokens at Flash pricing.
- GLM 5.3 Flash. Z.ai’s 320B-parameter MoE with 18B active, natively multimodal, 1M context, MIT.
- Qwen 3.8 Flash-Next. 125B parameters with only 6B active, plus a 51B n-gram table that can sit in system RAM. An early preview of the Qwen 4 architecture, under qwen-community rather than MIT.
- Gemini 3.8 Flash. Google’s third Flash in six weeks, and on DeepSWE v1.1 it beats most larger closed models. It holds 3.7’s $0.75 and $3.75 per million tokens, but only until December 31.
- Qwen 3.8 Max 0902. A new snapshot of Qwen’s flagship at the same price: 2.4T parameters, 1M context, and TerminalBench 3.0 up from 11.3 to 29.0.
- Muse Spark 1.3. Meta’s new flagship, closed for now with open weights promised. Two SKUs of the same model: $1.25 and $4.25 per million tokens, or $0.10 and $0.20 if you let Meta train on your traffic.
- OpenAI GPT-6 Astra. It shipped on September 3 at $10 and $50 per million tokens, and it is the first model OpenAI rates “critical” on cyber, at 100% on ExploitBench.
Which is the other pattern worth noticing. Three labs now have a cyber-capable model and each one picked a different gate: Z.ai post-trained GLM 5.3 for security work and published the weights, Google ships Gemini 3.8 Flash Cyber only to defenders it has vetted through its Fairwind programme, and OpenAI ships Astra broadly but keeps its cyber capabilities for approved defenders, through Daybreak Blue.
The frontier isn’t only moving toward bigger models anymore. It’s moving toward models that are good enough to run everywhere, all the time, for much less.
// from our side
One trillion tokens. That’s how many have now gone through Helmcode, and not generated in a benchmark or promised on a pricing page: served. We marked it the way we know how, by publishing what we’ve learned.
AI for local governments: a real deployment story
The flash wars run through our racks too. GLM 5.3 and DeepSeek v4 Flash Vision Exp are live on Enterprise, and GLM 5.3 replaces 5.2 on every Enterprise account, with the full million tokens of context served and at the same price.
GLM 5.3 is now available on Helmcode
Limits on NaN go up as well: GLM 5.3 Flash from 500M to 2B tokens a month, DeepSeek v4 Flash from 2B to 3B and now on the Vision Exp build, and the €200 tier adds GLM 5.3 itself at 3B tokens a month.
And two dates worth saving, both in Madrid:
- September 22, Café Helmcode. Small, informal, coffee and open models. We said September 23 in issue one; our mistake.
- October 6, Helmconf. The event we teased as Inference_ #01 has a name now, and Hugging Face, Cloudflare and InditexTech are already on board.
Seven models in two weeks. Let’s see how many we have to cover in issue three.
Borja