Docs

Rate limits

Rate limits apply per API key — RPM and concurrency — not to your total token volume.

Per-key limits

LimitValue
Requests per minute100 rpm
Concurrency5 concurrent requests per model — 10 on GLM 5.3, GLM 5.3 Flash, GLM 5.2, DeepSeek V4 Flash and Qwen 3.8 Flash
Tokens per minute2M tpm (language models)

Consolidated keys multiply every limit above by the number of seats on the key.

Embedding and speech models have their own request limits. See Models. GLM 5.3 is an add-on with its own limits and its own allowance, further down this page.

Monthly volume

Efficient open models are unlimited on every plan. The frontier-class models — DeepSeek V4 Flash and GLM 5.3 Flash — count against a monthly quota that scales with your plan:

PlanMonthly quota
Starter5B tokens
Growth15B tokens
Scale35B tokens
EnterpriseCustom

GLM 5.3 (the add-on) carries its own separate allowance — 2B / 6.5B / 18B tokens by plan — further down this page; the two quotas never draw from each other.

If you exceed your quota, we contact you before charging for the overage. See Pricing for the full breakdown.

Images

Image generation has its own monthly allowance, per organisation, and it does not draw from the token quota above:

PlanMonthly images
Starter300
Growth600
Scale1,200
EnterpriseCustom

Two things decide how fast that runs down:

  • The cap counts images produced, not requests. One call asking for n: 4 spends four.
  • Deleting an image does not refund it. The month’s usage is what was generated, not what you kept.

One bucket per organisation, shared by every surface: images made in the console and images made through /v1/images/* come out of the same total.

GLM 5.3

GLM 5.3 is a per-key add-on at €150 a month, and no plan covers it. It carries its own per-key limits:

LimitValue
Requests per minute100 rpm
Concurrency10 concurrent requests
Tokens per minute3M tpm
Context1M per request

Its monthly allowance belongs to your organisation and is shared across every key with GLM 5.3. It is separate from the quota above and never draws from it:

PlanGLM 5.3 tokens a month
Starter2B tokens
Growth6.5B tokens
Scale18B tokens

When the allowance runs out, GLM 5.3 pauses until the next billing cycle and everything your plan covers keeps answering. There are no overage charges here: you cannot spend more than you contracted. See Models.

Handling 429

When you exceed a limit, the API returns 429 Too Many Requests. Back off and retry. Most OpenAI SDKs already do this for you with exponential backoff — the Retry-After header tells you how long to wait.

# the OpenAI SDK retries 429s automatically — tune the count if needed
client = OpenAI(
    base_url="https://api.helmcode.com/v1",
    api_key="sk-your-key-here",
    max_retries=5,
)

Need higher limits?

High-demand workloads can get higher RPM and pooled concurrency on request. Dedicated and Enterprise plans also run on hardware reserved for you, with guaranteed throughput. Talk to us.