// opendata

The real numbers
behind the platform.

Real usage data from our inference platform, aggregated and anonymized. No prompts, no content; just request-level counters.

Updated weekly · 30 Aug 2026

1.1T

Tokens processed

cumulative

96.9M

Requests served

cumulative

20

Active models

in production

// models

Tokens by model.

Cumulative tokens processed per model. One open model carries most of the load, and the full stack is always available.

01 Qwen 3.6 48.8% 528.6 B · 82.6 M
02 DeepSeek V4-Flash 21.6% 233.9 B · 4.2 M
03 MiMo V2.5 14.2% 154.1 B · 2.2 M
04 deepseek-v4-flash-0731 5.5% 59.6 B · 512.6 K
05 glm5.2 5% 54.1 B · 617.6 K
06 Gemma 4 1.9% 20.6 B · 4.2 M
07 qwen3.8-flash 1.4% 14.9 B · 277.4 K
08 glm5.3-flash 1% 10.6 B · 96 K
09 Qwen3 Embedding 0.6% 6.9 B · 1.8 M
10 Qwen3 Coder <0.1% 142.8 M · 7.1 K
11 claude-opus-5 <0.1% 1.7 M · 533
12 claude-fable-5 <0.1% 2.2 K · 3
13 deepseek <0.1% 103 · 2
14 claude-haiku-4-5 <0.1% 43 · 3
15 Whisper Large v3 <0.1% 24 · 254.3 K
16 gpt-5.6-luna <0.1% 15 · 1
17 claude-sonnet-5 <0.1% 9 · 1
18 gemini-3.5-flash-lite <0.1% 7 · 1

// tokens

Input vs. output.

Inference here is overwhelmingly read-heavy, long prompts, retrieval and context, with a thin slice of generated tokens.

Input · prompt 1.1 T 97.8%
Output · generated 23.3 B 2.2%

// usage

Tokens per day.

Daily tokens processed over the last 90 days, peaking at 20.2 B/day.

0 6 B 12 B 18 B 24 B 20.2 B peak
90 days agotoday

// beyond text

Speech & reranking.

The stack is more than LLMs, transcription, synthesis and reranking run on the same API.

Text-to-speech Kokoro 92.3 K requests
Reranking Qwen3 Reranker 50.5 K requests

// clients · last 7 days

How teams connect.

Drop-in OpenAI compatibility in the wild, the official SDK and OpenCode account for the vast majority of traffic.

// geography · last 7 days

Where requests come from.

70.1% of traffic originates inside the EU, the audience this infrastructure is built for.

Spain 37.6% 2.6 M
AT 12.9% 885.8 K
Colombia 10.7% 739 K
Germany 9.9% 677.8 K
Mexico 7.9% 542.8 K
United States 6.2% 426.6 K
France 4.5% 310.4 K
Finland 3.9% 266.7 K
Argentina 2.9% 200.9 K
Ireland 0.9% 58.8 K
Netherlands 0.6% 37.9 K
Chile 0.4% 30.5 K
Others 1.7% 117 K

// performance

Latency & throughput.

Median time to first token and sustained throughput per model, measured on 30 Aug 2026.

Model TTFT p50 Throughput
Qwen 3.6 n/a 30648 rpm
DeepSeek V4-Flash n/a 29989 rpm
MiMo V2.5 n/a 27761 rpm
Gemma 4 n/a 27743 rpm
glm5.3-flash n/a 3712 rpm
qwen3.8-flash n/a 1335 rpm
Kokoro n/a 47.7 rpm
Whisper Large v3 n/a 30.8 rpm
Qwen3 Embedding n/a 14.2 rpm
Qwen3 Reranker n/a 12.9 rpm
glm5.2 n/a 9.2 rpm
claude-opus-5 n/a n/a

TTFT p50 = median time to first token · Throughput = sustained requests per minute.

// who it's for

Two ways to run on this stack.

These numbers come from real workloads across the community and private deployments alike.

Builders & community

Frontier models, fair price, no data sharing.

Access the latest open models at a reasonable cost, without handing over your data, through the NaN community.

nan.builders →
Startups & enterprise

Private, dedicated inference with SLAs.

Dedicated infrastructure, support and contractual SLAs, flat rate, EU data, OpenAI-compatible.

see_pricing →

Methodology. Figures are aggregated, anonymized counters collected at the request level. Helmcode keeps zero logs, no prompt or completion content is ever stored. Cumulative metrics span the platform's lifetime; windowed metrics are labelled per section.

// get started

START BURNING TOKENS

Skip the AI infra work. Deploy your first private inference endpoint today.

Flat rate. EU data. OpenAI API compatible.