// pricing

Pricing

Pay per API key, not per token. Unlimited tokens on Qwen and Gemma, no lock-in, cancel anytime.

Starter

€399/mo

For small teams starting to integrate AI into their workflows.


  • 5 API Keys
  • Unlimited tokens · Qwen and Gemma
  • 2.5B tokens/month · SOTA and DeepSeek
  • OpenAI API compatible
  • Zero logs · EU data
  • Email support

Scale

€3,199/mo

For organizations with intensive AI use and advanced needs.


  • 40 API Keys
  • Unlimited tokens · Qwen and Gemma
  • 20B tokens/month · SOTA and DeepSeek
  • OpenAI API compatible
  • Zero logs · EU data
  • Priority support
  • SLA 99.9%
  • Early access to new models

Enterprise

Custom

For organizations needing dedicated GPUs and custom configuration.


  • +60 API Keys
  • Unlimited tokens · Qwen and Gemma
  • Custom cap · SOTA and DeepSeek
  • Dedicated GPUs
  • Custom models
  • Zero logs · EU data
  • Custom SLA
  • Dedicated onboarding

All plans include RPM limits and concurrency per API Key to guarantee service quality.

// roi calculator

How much you save by migrating.

Compare what you spend on closed APIs with Helmcode's flat rate on open models. Three questions, no digging through token invoices.

€/month

list price $5 / $30 per M

what do you use it for? (pick as many as apply)

At least one use case stays selected.

Answer the three questions and press calculate_saving to see your estimate.

Indicative estimate for internal discussion. It starts from list prices per million tokens (July 2026) in US dollars, converted to euros at 0.867 (ECB reference rate, 2026-08-13), and from typical cache rates per use case and provider. We do not model cache write surcharges, so the estimate errs low rather than high. Helmcode does not guarantee a specific saving: your volume discount or your usage pattern can move the figure.

dedicated infrastructure

Your own inference stack,
in your datacenter.

If your use case requires total data sovereignty, we deploy and operate the full inference stack inside your own infrastructure. Your models, your data, your prompts, they never leave your network.

talk_to_us →
On-premise deployment We install and operate the full stack on your servers, with the same OpenAI-compatible API.
Hardware advisory We help you select the right GPUs, memory and network for your use case and budget.
Total data sovereignty Your data and prompts never leave your network. Built for banking, healthcare and defense.
In production at

// pricing faq

Pricing, explained.

Everything about plans, limits and billing, before you ask.

Do I pay per API key or per token?

Per API key, a flat monthly price. Qwen and Gemma carry no per-token charge at all, and everything else runs against an allowance you know in advance. Your CFO gets a fixed line on the P&L.

What's the difference between "unlimited tokens" and the monthly SOTA cap?

Qwen3.6 and Gemma 4 are unlimited on every plan, along with speech and embeddings. The monthly cap covers the models where compute is more expensive, and DeepSeek V4 Flash counts against it. GLM 5.2 has its own separate allowance and never touches your cap. We always reach out before any overage, never a surprise bill.

How is GLM 5.2 billed?

As an add-on to any plan, €150 per API key per month. It comes with its own monthly token allowance for your organisation, shared across your keys: 2B on Starter, 6.5B on Growth, 18B on Scale. That allowance is separate from the SOTA cap, so heavy GLM use never eats the rest of your capacity, and there are no overage charges: usage pauses until the next billing cycle.

Is there any commitment or lock-in?

No. Plans are month-to-month and you can cancel anytime. You run on open-weight models you can always access, no vendor can deprecate your API or change pricing on you overnight.

Can I change plans later?

Yes. Upgrade or downgrade at any time and changes are prorated. As your usage grows you simply move up a tier, the API and your code stay exactly the same.

How do rate limits work?

Limits apply per API key as requests-per-minute and concurrency, to guarantee service quality, not on how many tokens you process. A single key can handle hundreds of millions of tokens a month.

Do you offer dedicated GPUs or on-premise?

Yes, on Enterprise: dedicated NVIDIA Blackwell hardware, custom and fine-tuned models, and full on-premise deployment inside your own datacenter. Talk to us for a custom quote.

// get started

START BURNING TOKENS

Skip the AI infra work. Deploy your first private inference endpoint today.

Flat rate. EU data. OpenAI API compatible.