// open models · european infrastructure

DeepSeek, GLM and Qwen,
served from Europe

The best open models, served on our infrastructure without DeepSeek, Z.ai or Alibaba seeing your data.

  • Own hardware in the EU
  • No prompt logs
  • GDPR DPA
  • OpenAI-compatible

// where your data is processed

Open models running
on our machines

DeepSeek, Z.ai and Alibaba publish the weights of these models and we run them on our own servers. The weights were downloaded once, when the model was deployed, and the developers receive no prompts, no responses and no usage data.

client Your app any OpenAI-compatible SDK
entry Helmcode API TLS 1.3 · API key
inference Helmcode GPU in memory · no logs open weights loaded DeepSeek V4.1 FlashGLM 5.3Qwen 3.6

// request lifecycle

What happens
to a request

The request enters our infrastructure, is processed in memory and the response returns to your application. We do not keep the content of the prompt or the response once inference ends.

01 Enters TLS · EU infrastructure
02 Is processed Model loaded on GPU · in memory
03 Responds Streaming to your application
04 Is discarded No prompt or response is kept

// models

The models we operate
in Europe

Open weights published by their developers, served from the same Helmcode API.

ModelDeveloperServed byRuns inTypeContextLicenceEndpoint
DeepSeek V4.1 Flash DeepSeek · China Helmcode Europe LLM 1M MIT deepseek-v4-flash
GLM 5.3 Z.ai · China Helmcode Europe multimodal LLM 1M GLM-5.3 License* glm5.3
GLM 5.3 Flash Z.ai · China Helmcode Europe multimodal LLM 1M MIT glm5.3-flash
Qwen 3.6 Alibaba · China Helmcode Europe multimodal LLM 256K Apache 2.0 qwen3.6
Qwen3 Embedding 8B Alibaba · China Helmcode Europe embeddings · Apache 2.0 qwen3-embedding
Qwen3 Reranker 8B Alibaba · China Helmcode Europe reranking · Apache 2.0 rerank

* GLM-5.3 License: MIT plus a Z.ai clause for anyone earning over $10bn a year serving the model.

// privacy

Your data stays
safe and private

Every request is processed in memory, in the EU, and discarded when inference ends. This is everything that happens to what you send to the open models we operate.

Processed · transient

Only while inference runs

The prompt and the response live in GPU memory for the length of the request, inside the EU.

  • Your prompt, in memory, in the EU
  • Inference on the model
  • The response, streamed back to you

Stored · for billing only

Just enough to bill you

Metadata and counts that let us operate and invoice the service. Never the content.

  • API key metadata
  • Aggregate request and token counts

Never stored

Nothing you send stays

No prompt or response is written to disk, and nothing you send trains a model.

  • Prompts
  • Responses
  • Your documents or code
  • Anything that trains a model

// cost

The difference shows
on the invoice too

Example for 10 billion tokens a month, 80% input and 20% output.

  • Claude Opus 5.5 via API

    list price

    ~$72,000/ month

  • DeepSeek V4.1 Flash via the developer’s API

    off-peak price; it doubles at peak hours

    ~$2,400/ month

  • Helmcode Growth

    includes 15B tokens a month of DeepSeek V4.1 Flash and GLM 5.3 Flash; Qwen 3.6 unlimited

    €1,299/ month

The comparison shows the cost of processing that volume of tokens, not an equivalence of quality or capability between the models.

Official prices in October 2026. Bars on a linear scale of each figure.

see_pricing →

// security and procurement

What your security team
will ask you

These answers cover the open models Helmcode operates. Copy them as they are, or request the DPA and the full pack.

  1. 01

    Is personal data transferred outside the EEA during inference?

    No, for the open models Helmcode operates: inference runs on Helmcode infrastructure in EU data centres. Closed models resold through the same API run on their providers’ infrastructure, under their terms.

  2. 02

    Are DeepSeek, Z.ai or Alibaba subprocessors?

    No. They publish the model weights. Helmcode downloads them and runs them on its own infrastructure; the developers receive no requests and no customer data.

  3. 03

    Are prompts or responses stored or logged?

    No. Prompts and responses are processed in memory and discarded when inference ends. Helmcode keeps only API key metadata and aggregate request and token counts, for operating and billing the service.

  4. 04

    Is customer data used to train or fine-tune models?

    No. Nothing sent to the API is used to train or fine-tune any model.

  5. 05

    Can the model send data to a third party by itself?

    No. The weights are loaded into the inference engine Helmcode operates. The model has no network connection of its own to DeepSeek, Z.ai, Alibaba or anyone else.

  6. 06

    Which jurisdiction does the service operate under?

    The EU. These models run on infrastructure Helmcode owns and operates in the EU, and Helmcode signs a Data Processing Agreement under Article 28 of the GDPR.

request_DPA_and_security_pack →

// faq

Details that usually
come up next

Is it exactly the same model the developer offers?

It is the open-weights release the developer published. Some models are served quantised to NVFP4 for efficient inference, so answers can differ slightly from the developer’s endpoint. Test it on your own cases before production.

Where are the servers?

In EU data centres, on hardware Helmcode owns and operates.

What about the closed models available on Helmcode?

They run on their providers’ infrastructure, under their terms. Everything on this page about EU processing and no international transfer refers only to the open models Helmcode operates.

Can I deploy it on-premise?

Yes. The on-premise deployment runs the same models and the same API inside your own infrastructure.

Try these models
on your own workload

The same OpenAI-compatible API. Inference in the EU, no prompt logs and a flat rate.