GLM 5.2 is now available on Helmcode

GLM 5.2 is now available on Helmcode

GLM 5.2, the 744B MIT open model, now runs on Helmcode's European infrastructure with zero logs and a flat rate of 150 euros a month.

The reference and most balanced open model, served on European infrastructure

On 17 June, Z.ai released the most capable open model published to that date under an MIT licence.

As of today, GLM 5.2 is available to every Helmcode Enterprise account.

It runs on our European infrastructure, with zero logs, at 150 euros a month for each API key you use with it. A flat rate with a monthly token allowance, so there are no surprises.

We have spent more than 3 weeks testing it, tuning it and using it before offering it. It has been available for two weeks inside the NaN community so we could test concurrency and load capacity.

Our summary: it is the best balanced open model on cost and benefit, and it matches Claude Opus 4.8 on agentic coding. It delivers enormous quality for development work and agent orchestration without the strings and the drawbacks that come with closed frontier models.

GLM 5.2 in one minute

GLM 5.2 is Z.ai 's flagship: a 744 billion parameter MoE, of which only around 40 billion activate per token. It ships with a 1 million token context and two reasoning modes, High and Max.

The two most innovative parts of this model:

  • IndexShare makes that 1 million token context genuinely usable. It reuses the same attention indexer as GLM 5.1, every four layers. That cuts compute per token by 2.9 times at full context.
  • MTP (Multi Token Prediction) improved by 20%. The model generates faster and consumes less for the same compute. The ablation was run on GLM 5.1, but the same technique has been applied to this new model.

You can read the rest of the architecture in Z.ai's launch post .

The important part is that all the weights were published under an MIT licence, which lets us run it and make it available on our own infrastructure, privately, for our customers. Nobody can take access to the model away. Nobody logs the model's prompts and outputs.

This is Helmcode's constant fight: that companies stay sovereign over their tools and their data.

Benchmarking GLM 5.2

Most of the benchmarks available in the first few days came from Z.ai , but there are 3 companies in particular that we give more credibility to (Proximal, PostTrainBench and Abundant AI) which measured the model at its full 1M token context and maximum effort.

  • On FrontierSWE (open technical projects running tens of hours) GLM 5.2 lands within 1% of Claude Opus 4.8 and beats GPT-5.5.
  • On PostTrainBench, second only to Opus 4.8, ahead of GPT-5.5.
  • And on SWE-Marathon, the ultra long horizon engineering benchmark (building compilers, optimising kernels), is where the biggest gap shows: a score of 13.0 against 26.0 for Opus 4.8. Half.

Outside third party measurements, there is one figure in Z.ai 's own table worth mentioning: on competition mathematics (AIME 2026) it reaches 99.2, above every closed model.

The real profile: put those results together with the fact that it is an open weights model, and what you have is a genuine machine for long context agentic coding and mathematics at a fraction of the price of the closed models.

It sits only a little below the frontier on general scientific reasoning (91.2 on GPQA-Diamond against 93 to 94 for the closed models, according to Z.ai 's table) and on the most extreme engineering tasks.

GLM 5.2 is not the best model for everything. But it matches, and sometimes beats, the frontier models on what matters most to our customers.

What about Kimi K3?

That is the obvious question. On 16 July, Moonshot launched Kimi K3 and it scores higher on Artificial Analysis' Intelligence Index: 57 against 51 for GLM 5.2. On the overall index, K3 is today the most capable open model. We are not going to argue with that figure, because transparency is part of the product too.

Now, the nuances, which matter a great deal here:

  • Intelligence is only one of the three indices. On the other two from Artificial Analysis the scoreboard flips: GLM 5.2 runs at 179 tokens per second against 35 for K3 (5 times faster) and costs 0.59 dollars per completed task against 0.86 (a third less). GLM 5.2 is faster and cheaper.
  • The licences are not comparable. GLM 5.2 is pure MIT: anyone can run it, serve it and build a business on top without asking permission. K3 shipped under its own licence, with commercial conditions that trigger on revenue. For an inference provider, or for a company that wants to self host with guarantees, that is a significant difference.
  • Size has consequences. K3 is 2.8 trillion parameters, almost 4 times GLM 5.2. That translates into more hardware and more cost per token (its API runs around 3 times more expensive on output). GLM 5.2 runs on a fraction of the hardware, and that margin is exactly what lets us offer it at this price on a flat rate.
  • K3 has native vision (images and video). GLM 5.2 is text only, although the Baseten team have added vision to it .

And on real agentic coding, which is our customers' central use case? There the practical difference is minimal. We measured it ourselves.

We tested it on our own infrastructure

We did not stop at other people's benchmarks. We ran Kimi K3 on our own hardware (4 B300 GPUs, quantised in NVFP4, tested on both vLLM and SGLang) and put it up against GLM 5.2 on the ground where K3 is supposed to shine: very long agentic tasks, hundreds of consecutive steps.

The test: two complete projects (frontend, backend and infra) run in parallel with teams of agents, one on each model, plus a battery of public and in house evals. The result:

  • Both finished the projects in around 5 hours with excellent results. Both followed the guidelines and implemented validations, tests and middleware. K3 finished about 15 minutes earlier and has slightly better taste in frontend. GLM 5.2 came out functionally level.
  • The jump K3's benchmarks promise does not show up in real projects. The practical difference between the two is marginal.
  • The difference that is enormous: compute. K3 needed nearly twice the compute of GLM 5.2 for the same throughput. For an inference provider, that means nearly twice the cost to serve. That difference is exactly what separates being able to offer a flat rate from not being able to.

At equal results on real projects, GLM 5.2 costs half as much to serve, has a clean MIT licence to run it on private infrastructure, and makes a flat, predictable price possible. On quality against price there is no contest. That is why it is our flagship.

In short: of the open models with frontier class intelligence (50 or above on the Artificial Analysis index) there are three today. We serve two of them, GLM 5.2 and DeepSeek V4 Flash, and GLM 5.2 is the fastest of the three: 179 tokens per second against 123 and 35.

The twelve models with frontier class intelligence compared across three panels. On intelligence, GLM 5.2 scores 51 and is the second of the open models, behind Kimi K3 on 57 and ahead of DeepSeek V4 Flash on 50. On speed, GLM 5.2 is second in the whole table at 179 tokens per second. On cost per task it comes in at 0.59 dollars. Source: Artificial Analysis, Intelligence Index v4.1, captured 3 August 2026.

We will keep publishing whatever we measure on K3. If its licence and its cost to serve ever fit our model, we will say so here with the same numbers in hand.

What do I gain going through Helmcode instead of Z.ai's own API?

This is the good part about open weights models. You get to use the best MIT licensed open model of the moment without having to send your data to Z.ai 's API.

With European models in short supply, open models solve the data sovereignty problem.

We serve it from our European infrastructure. We store no logs, so your information stays yours alone. Not a single token leaves our machines.

With AI Act penalties applying since 2 August, where and how your inference is served has stopped being a technical detail.

The alternative is running it yourself, on your own machines. Which is also something we can do. If you have machines available with around 860 GB of VRAM (an 8x H200 node for the FP8 checkpoint), go ahead. We can operate the infrastructure for you.

The price

Making a model this size available has a consequence: we cannot sell it at the same price as the rest of our models. The infrastructure it needs is far more expensive and the margins are thinner.

So rather than raise the price of every plan, we decided that only those who want to use it can add it to theirs. Add as many API keys with GLM 5.2 as you need to your plan, and that is it.

Each API key with GLM 5.2 is 150 euros a month as an add-on to any enterprise plan.

Each key with GLM 5.2 includes: 100 requests per minute, 5 concurrent processes, 3M tokens per minute and context up to 500K per request. The model takes 1M of context and we serve 500K on purpose: it is what lets us keep concurrency high and hold the price at 150 euros.

How much those keys can consume in a month comes from your plan tier, as a token allowance shared across every key on the account:

PlanTokens per monthDaily average over a working month
Starter2,000Mabout 95M a day
Growth6,500Mabout 310M a day
Scale18,000Mabout 860M a day

When the allowance runs out, usage pauses until the next billing cycle. There are no overage charges: it is impossible to spend more than you contracted.

When does it pay off against paying per token?

We ran the numbers (assumption: 3:1 input to output ratio, July 2026 prices):

Monthly usageOpenRouter (about €0.80/M)Z.ai API (about €2.00/M)Helmcode (flat rate)
100M tokens€80€200€150, not yet worth it here
400M tokens€320€800€150
2,000M tokens (the Starter allowance)€1,600€4,000€150
6,500M tokens (the Growth allowance)€5,200€13,000€150
18,000M tokens (the Scale allowance)€14,400€36,000€150

The simple answer: past roughly 190M tokens a month the rate pays for itself, which is the point where OpenRouter reaches those same 150 euros. Everything beyond that is margin. A single key burning through the Starter allowance pays an effective €0.075/M, and if that key eats a whole Scale allowance the effective price drops below €0.01/M. No per token API reaches those numbers.

And the answer we would give you if you asked: if your volume is low and your data can travel to an American router, OpenRouter will come out cheaper. This product is for whoever consumes volume, needs the data in Europe, or needs to budget without shocks.

(Note for the NaN community: GLM 5.2 will arrive as its own 200 euros a month tier in the coming days. The tier will keep access to the rest of the models, with GLM 5.2 added.)

How to turn it on

It is available right now on the starter, growth and scale plans.

  • If you're already a Helmcode customer, go to your dashboard and generate a new SoTA API key. You can assign it to GLM 5.2, and the usage cost will be added to your next invoice.
  • You can also go to Billing → SoTA Models to enable SoTA models either for your entire team or only for specific API keys.
  • If you are not a customer yet but want to be, when you pick your plan (starter, growth or scale) the next step will ask how many API keys with GLM 5.2 you need.

After that, operating it is simple. The endpoint is OpenAI compatible:

from openai import OpenAI

client = OpenAI(base_url="https://api.helmcode.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "..."}],
)

Full docs here: helmcode.com/docs

We have put the most balanced model on the market at the best price within reach of every company. Now it is your turn to squeeze it. Start burning tokens :)

undefined

The Helmcode digest: open models, releases, what is happening in open AI, opinions and common sense. Published twice a month.