// opendata

The real numbers
behind the platform.

Real usage data from our inference platform, aggregated and anonymized. No prompts, no content; just request-level counters.

Updated weekly · 15 Jul 2026

541.5B

Tokens processed

cumulative

52.5M

Requests served

cumulative

10

Active models

in production

// models

Tokens by model.

Cumulative tokens processed per model. One open model carries most of the load, and the full stack is always available.

01 Qwen 3.6 61.1% 330.7 B · 46.2 M
02 MiMo V2.5 15.1% 81.7 B · 1.1 M
03 DeepSeek V4-Flash 14.6% 79.1 B · 1.6 M
04 glm5.2 6.3% 34.2 B · 376.3 K
05 Gemma 4 2.3% 12.5 B · 2.3 M
06 Qwen3 Embedding 0.6% 3.1 B · 870.5 K
07 Qwen3 Coder <0.1% 142.8 M · 7.1 K
08 Whisper Large v3 <0.1% 24 · 70.6 K

// tokens

Input vs. output.

Inference here is overwhelmingly read-heavy, long prompts, retrieval and context, with a thin slice of generated tokens.

Input · prompt 529.5 B 97.8%
Output · generated 12 B 2.2%

// usage

Tokens per day.

Daily tokens processed over the last 90 days, peaking at 13.9 B/day.

0 4 B 8 B 12 B 16 B 13.9 B peak
90 days agotoday

// beyond text

Speech & reranking.

The stack is more than LLMs, transcription, synthesis and reranking run on the same API.

Text-to-speech Kokoro 35.6 K requests
Reranking Qwen3 Reranker 20.3 K requests

// clients · last 7 days

How teams connect.

Drop-in OpenAI compatibility in the wild, the official SDK and OpenCode account for the vast majority of traffic.

OpenAI SDK (Python) 39.5% 236 devs
OpenAI SDK (JS) 13.1% 78 devs
curl 12.2% 73 devs
Node.js 8.4% 50 devs
Vercel AI SDK 8.2% 49 devs
Python httpx 6% 36 devs
Go HTTP client 4.2% 25 devs
Python requests 3.9% 23 devs
Others 4.5% 27 devs

// geography · last 7 days

Where requests come from.

75.2% of traffic originates inside the EU, the audience this infrastructure is built for.

Spain 30.2% 950.8 K
France 17.7% 558.3 K
Germany 14.7% 461.7 K
Colombia 13.6% 429.3 K
Finland 12.1% 382.5 K
United States 5% 158 K
United Kingdom 1.9% 60.6 K
Argentina 1.7% 52.4 K
Mexico 1.3% 42.1 K
Chile 0.5% 16.5 K
Ireland 0.2% 7.3 K
Belgium 0.2% 7 K
Others 0.7% 22.6 K

// performance

Latency & throughput.

Median time to first token and sustained throughput per model, measured on 15 Jul 2026.

Model TTFT p50 Throughput
Qwen 3.6 3.4 s 4708090 rpm
DeepSeek V4-Flash 7.8 s 12099 rpm
Gemma 4 177 ms 2509 rpm
MiMo V2.5 2.1 s 1777 rpm
Qwen3 Embedding n/a 1403 rpm
Whisper Large v3 n/a 16.6 rpm
Kokoro n/a 1 rpm
Qwen3 Reranker n/a 0.5 rpm
glm5.2 n/a 0.4 rpm

TTFT p50 = median time to first token · Throughput = sustained requests per minute.

// who it's for

Two ways to run on this stack.

These numbers come from real workloads across the community and private deployments alike.

Builders & community

Frontier models, fair price, no data sharing.

Access the latest open models at a reasonable cost, without handing over your data, through the NaN community.

nan.builders →
Startups & enterprise

Private, dedicated inference with SLAs.

Dedicated infrastructure, support and contractual SLAs, flat rate, EU data, OpenAI-compatible.

see_pricing →

Methodology. Figures are aggregated, anonymized counters collected at the request level. Helmcode keeps zero logs, no prompt or completion content is ever stored. Cumulative metrics span the platform's lifetime; windowed metrics are labelled per section.

// get started

START BURNING TOKENS

Skip the AI infra work. Deploy your first private inference endpoint today.

Flat rate. EU data. OpenAI API compatible.