Rate limits
Rate limits apply per API key — RPM and concurrency — not to your total token volume.
Per-key limits
| Limit | Value |
|---|---|
| Requests per minute | 100 rpm |
| Concurrency | 5 concurrent requests per model — 10 on GLM 5.3, GLM 5.3 Flash, GLM 5.2, DeepSeek V4 Flash and Qwen 3.8 Flash |
| Tokens per minute | 2M tpm (language models) |
Consolidated keys multiply every limit above by the number of seats on the key.
Embedding and speech models have their own request limits. See Models. GLM 5.3 is an add-on with its own limits and its own allowance, further down this page.
Monthly volume
Efficient open models are unlimited on every plan. The frontier-class models — DeepSeek V4 Flash and GLM 5.3 Flash — count against a monthly quota that scales with your plan:
| Plan | Monthly quota |
|---|---|
| Starter | 5B tokens |
| Growth | 15B tokens |
| Scale | 35B tokens |
| Enterprise | Custom |
GLM 5.3 (the add-on) carries its own separate allowance — 2B / 6.5B / 18B tokens by plan — further down this page; the two quotas never draw from each other.
If you exceed your quota, we contact you before charging for the overage. See Pricing for the full breakdown.
Images
Image generation has its own monthly allowance, per organisation, and it does not draw from the token quota above:
| Plan | Monthly images |
|---|---|
| Starter | 300 |
| Growth | 600 |
| Scale | 1,200 |
| Enterprise | Custom |
Two things decide how fast that runs down:
- The cap counts images produced, not requests. One call asking for
n: 4spends four. - Deleting an image does not refund it. The month’s usage is what was generated, not what you kept.
One bucket per organisation, shared by every surface: images made in the console and images made through /v1/images/* come out of the same total.
GLM 5.3
GLM 5.3 is a per-key add-on at €150 a month, and no plan covers it. It carries its own per-key limits:
| Limit | Value |
|---|---|
| Requests per minute | 100 rpm |
| Concurrency | 10 concurrent requests |
| Tokens per minute | 3M tpm |
| Context | 1M per request |
Its monthly allowance belongs to your organisation and is shared across every key with GLM 5.3. It is separate from the quota above and never draws from it:
| Plan | GLM 5.3 tokens a month |
|---|---|
| Starter | 2B tokens |
| Growth | 6.5B tokens |
| Scale | 18B tokens |
When the allowance runs out, GLM 5.3 pauses until the next billing cycle and everything your plan covers keeps answering. There are no overage charges here: you cannot spend more than you contracted. See Models.
Handling 429
When you exceed a limit, the API returns 429 Too Many Requests. Back off and retry. Most OpenAI SDKs already do this for you with exponential backoff — the Retry-After header tells you how long to wait.
# the OpenAI SDK retries 429s automatically — tune the count if needed
client = OpenAI(
base_url="https://api.helmcode.com/v1",
api_key="sk-your-key-here",
max_retries=5,
)
Need higher limits?
High-demand workloads can get higher RPM and pooled concurrency on request. Dedicated and Enterprise plans also run on hardware reserved for you, with guaranteed throughput. Talk to us.