// opendata

The real numbers
behind the platform.

Real usage data from our inference platform, aggregated and anonymized. No prompts, no content; just request-level counters.

Updated weekly · 9 Aug 2026

788.5B

Tokens processed

cumulative

76M

Requests served

cumulative

15

Active models

in production

// models

Tokens by model.

Cumulative tokens processed per model. One open model carries most of the load, and the full stack is always available.

01 Qwen 3.6 56% 441.4 B · 66.1 M
02 MiMo V2.5 16.4% 129 B · 1.7 M
03 DeepSeek V4-Flash 15.7% 123.8 B · 2.6 M
04 deepseek-v4-flash-0731 4.8% 37.7 B · 273.7 K
05 glm5.2 4.7% 36.9 B · 405.3 K
06 Gemma 4 2% 15.6 B · 3.3 M
07 Qwen3 Embedding 0.5% 3.9 B · 1.3 M
08 Qwen3 Coder <0.1% 142.8 M · 7.1 K
09 claude-fable-5 <0.1% 2.2 K · 3
10 claude-haiku-4-5 <0.1% 34 · 2
11 Whisper Large v3 <0.1% 24 · 204.9 K
12 gpt-5.6-luna <0.1% 15 · 1
13 gemini-3.5-flash-lite <0.1% 7 · 1

// tokens

Input vs. output.

Inference here is overwhelmingly read-heavy, long prompts, retrieval and context, with a thin slice of generated tokens.

Input · prompt 771.4 B 97.8%
Output · generated 17.1 B 2.2%

// usage

Tokens per day.

Daily tokens processed over the last 90 days, peaking at 15.9 B/day.

0 5 B 10 B 15 B 15.9 B peak
90 days agotoday

// beyond text

Speech & reranking.

The stack is more than LLMs, transcription, synthesis and reranking run on the same API.

Text-to-speech Kokoro 59.6 K requests
Reranking Qwen3 Reranker 40.1 K requests

// clients · last 7 days

How teams connect.

Drop-in OpenAI compatibility in the wild, the official SDK and OpenCode account for the vast majority of traffic.

OpenAI SDK (Python) 32.5% 196 devs
curl 17.9% 108 devs
OpenAI SDK (JS) 16.3% 98 devs
Node.js 8.8% 53 devs
Python httpx 5.8% 35 devs
Vercel AI SDK 5.8% 35 devs
Go HTTP client 5.5% 33 devs
Python requests 4% 24 devs
Others 3.5% 21 devs

// geography · last 7 days

Where requests come from.

71.3% of traffic originates inside the EU, the audience this infrastructure is built for.

Spain 51.6% 2.8 M
Colombia 14.4% 795.8 K
Germany 9.8% 537.5 K
United States 8.1% 444.2 K
Finland 6% 328.1 K
France 3.3% 179.9 K
Mexico 2.8% 156.9 K
Argentina 1.5% 83.3 K
Netherlands 0.5% 27.9 K
Chile 0.4% 21.9 K
Peru 0.3% 15.6 K
Ireland 0.2% 11.4 K
Others 1.2% 65 K

// performance

Latency & throughput.

Median time to first token and sustained throughput per model, measured on 9 Aug 2026.

Model TTFT p50 Throughput
Qwen 3.6 n/a 15611 rpm
DeepSeek V4-Flash n/a 12992 rpm
Gemma 4 n/a 12963 rpm
MiMo V2.5 n/a 12934 rpm
Whisper Large v3 n/a 16.3 rpm
Kokoro n/a 16.3 rpm
deepseek-v4-flash-0731 n/a 13.9 rpm
Qwen3 Embedding n/a 10 rpm
glm5.2 n/a 3.9 rpm
Qwen3 Reranker n/a 1.1 rpm

TTFT p50 = median time to first token · Throughput = sustained requests per minute.

// who it's for

Two ways to run on this stack.

These numbers come from real workloads across the community and private deployments alike.

Builders & community

Frontier models, fair price, no data sharing.

Access the latest open models at a reasonable cost, without handing over your data, through the NaN community.

nan.builders →
Startups & enterprise

Private, dedicated inference with SLAs.

Dedicated infrastructure, support and contractual SLAs, flat rate, EU data, OpenAI-compatible.

see_pricing →

Methodology. Figures are aggregated, anonymized counters collected at the request level. Helmcode keeps zero logs, no prompt or completion content is ever stored. Cumulative metrics span the platform's lifetime; windowed metrics are labelled per section.

// get started

START BURNING TOKENS

Skip the AI infra work. Deploy your first private inference endpoint today.

Flat rate. EU data. OpenAI API compatible.