undefined nº 03

Use DeepSeek without sending a single token to China

Hey,

There’s a common misunderstanding whenever DeepSeek, Qwen or GLM come up, which is that using a Chinese model means using a Chinese company’s API. It doesn’t: DeepSeek, Qwen and GLM publish their weights, so the models can be downloaded, deployed on someone else’s infrastructure and run without the lab that trained them taking part in a single request.

That is exactly what we do at Helmcode. We run some of the best open models on our own servers in data centres in Finland, Germany and France, so when you call DeepSeek V4.1 Flash, GLM 5.3 or Qwen through Helmcode, the request reaches our infrastructure, the model processes it there and the answer goes straight back to your application.

DeepSeek doesn’t get the prompt. Alibaba doesn’t get the prompt. Z.ai doesn’t get the prompt. And we don’t keep it either once the request is done.

// the weights may come from china, your data doesn’t have to go there

This matters a great deal when you’re trying to bring these models into a European company. If you use DeepSeek’s API directly, its own privacy policy says your data is processed and stored in China, which means your security and data protection people have to assess an international transfer before anything else.

With the models we operate, there is no such transfer for the inference, because the weights are already on our servers and the requests are processed inside the EU. Nor is the manufacturer a sub-processor, since it never receives any customer data.

None of this means that running a model in Europe takes care of your product’s GDPR by itself, and you still have to justify what data you process, for what purpose and on what legal basis. It does remove a big part of the problem, though: the data you send to the model doesn’t have to leave Europe.

DeepSeek’s privacy policy

// open models turn out to be much cheaper, too

Privacy would be reason enough to do it this way, but there’s another fairly obvious one, which is cost. Take a large but by no means unusual volume for a company starting to put AI into production, 10 billion tokens a month with 80% input and 20% output, at current list prices:

  • Claude Opus 5.5 through its API: about $72,000 a month.
  • DeepSeek V4.1 Flash on DeepSeek’s own API: about $2,400 a month at its off-peak rate, and twice that at peak hours.
  • Helmcode’s Growth plan: €1,299 a month, with 15 billion tokens of DeepSeek V4.1 Flash and GLM 5.3 Flash included.

So at this volume a flat plan in Europe comes in below even the manufacturer’s off-peak price. If you only need a fraction of those tokens, DeepSeek’s own API is still the cheapest way to run the model, as long as you can live with where it runs.

We aren’t comparing exactly the same model here, and we’re not saying DeepSeek replaces Opus for every task. The interesting part is something else: for the last two years, plenty of companies have assumed that putting a capable model into production meant accepting two things, paying per token to an American provider and sending it all the data the inference needs. Open models change that equation quite a lot, because you can have DeepSeek, GLM or Qwen behind an ordinary API without buying GPUs, building an infrastructure team or having every request travel to the model’s manufacturer.

// so can i put deepseek into production tomorrow?

That depends on what your application does, obviously, and our advice is exactly the same for DeepSeek as for Claude, GPT or Gemini: test the model on your own cases before you pick it.

The weights are still the ones the manufacturer published, so the model’s training, its biases and the way it answers on certain topics don’t change because we run it in Finland, Germany or France. We also serve some models quantised to make inference more efficient, which means an answer can differ slightly from what the original API would give you.

What has changed is that you now have another option, and you can judge DeepSeek on what it does as a model without having to accept how DeepSeek wants to handle your data.

Open models on European infrastructure

Pricing

// a few more things from this month

  • Claude Opus 5.5. Anthropic launched it at $4 per million input tokens and $20 per million output, down from $5 and $25.
  • MiMo-V2.6 Pro. Xiaomi published the training code under MIT along with the weights, which is still pretty rare.
  • Qwen3.8-Omni-Flash. Text, image, audio and video with a million tokens of context, although this time only through the API.
  • GLM-5.3. NIST’s CAISI rates it the most cyber-capable open-weight model it has evaluated so far.
  • Clef and Clef-flash. Cloudflare’s new decision models don’t write text: they return a probability for each allowed answer to a set of typed questions. Built on Qwen, with Apache 2.0 weights.
  • Mistral Large 4. A 1-trillion-parameter MoE trained from scratch in Mistral’s own European data centres, API-only for now, with the open weights promised by the end of October.

// from our side

We’ve added MiMo-V2.6 Flash to nan.builders, and Qwen-Image 2.1 to both nan.builders and Helmcode.

On Tuesday, September 22 we held our first Café Helmcode in Madrid, and it won’t be the last.

A short recap

Three days ago, on October 6, we were back in Madrid for Kernel Panic!, where Hugging Face, Cloudflare, InditexTech and Helmcode talked about what happens when these models leave testing and start going into production. The tickets were gone well before the day, so if you couldn’t make it, Cristian has written up how it went:

Kernel Panic!, by Cristian Cordova

Until the next one,

Borja

All issues