industries · Developer tools

Open coding models on a flat rate. Your margin stops depending on someone else’s pricing.

Inference is the COGS of a coding product: every point of gross margin rides on the price per token. Open coding models on a flat rate turn that variable into a fixed line.

the economics

The economics, by design.

Each cost lever mapped to a platform capability. Sovereignty comes by default.

Margin

inference = your COGS

The problemEvery point of gross margin depends on a third party’s per-token price.

HelmcodeA flat rate per API key: cost stops scaling with usage and your margin stops depending on external pricing.

Frontier code

GLM-5.2 · DeepSeek V4 Pro

The problemYou need frontier-grade code without frontier-API bills.

HelmcodeGLM-5.2 leads agentic coding (62.1% SWE-bench Pro); DeepSeek V4 Pro leads bounded generation (80.6% Verified). Both open, both self-hostable.

Sovereignty

EU · on-premise

The problemYour customers’ code should not travel to an external API.

HelmcodeOpen weights on EU infrastructure or on-premise: code never leaves your perimeter.

how coding ability is measured

What SWE-bench actually asks of a model.

Every vendor quotes a coding score at you. It is worth knowing what the benchmark behind those numbers does, because it is harsher than a puzzle set and because the numbers move fast enough that any figure without a date is worthless.

01

Real issues, real repositories

The benchmark is built from 2,294 software engineering problems taken from real GitHub issues and their corresponding pull requests across 12 popular Python repositories. The model gets a codebase and an issue description, and has to edit the code.

02

It spans files, not functions

The authors are explicit that resolving these issues frequently requires understanding and coordinating changes across multiple functions, classes and even files at once, interacting with execution environments and processing very long contexts. That is why context window and price per token stop being separable concerns.

03

Always ask for the date

When the benchmark was published in 2023, the best model of the day solved 1.96% of the issues. Current systems are far past that, which is exactly the point: a coding score without a date and a system version attached tells you nothing, ours included.

SWE-bench · Princeton NLP and collaborators "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?", arXiv:2310.06770. The official leaderboards at swebench.com separate open-source from proprietary systems and cover several variants, among them SWE-bench Verified and SWE-bench Multilingual, so a claim can be checked rather than trusted. read the report →

use cases

Your most common use cases.

The cases with the most traction in the sector, each with its own page in detail.

Recommended open models.

A starting point per task type. The full guide maps 80 cases to the open model for each one.

GLM-5.2MIT · 1M ctx
The open leader in agentic coding and long multi-step runs over real repos.
DeepSeek V4 ProMIT · 1M ctx · 1.6T/49B act
The king of bounded generation: 80.6% SWE-bench Verified, 93.5% LiveCodeBench.
DeepSeek V4 FlashMIT · 1M ctx in Helmcode
The flat-rate workhorse for high-volume, bounded coding tasks.

in progressWe are distilling and quantizing these open models into small, tightly specialised versions, trained for one task rather than for all of them. A model like that runs on less hardware, answers faster and fits where the big one does not, your own datacenter included. If you have a process with volume and stable criteria, that is the conversation we want to have with you.

// faq

Questions, answered.

What the sector's technical, compliance and business teams ask.

Why does inference cost decide our margin?

In a coding product, inference is the cost of goods sold: it scales directly with usage. On a per-token API, growth erodes gross margin. A flat rate per API key fixes the line, so scaling users no longer scales the bill.

How good is open code today?

GLM-5.2 leads agentic coding (62.1% SWE-bench Pro, 81 Terminal-Bench) and DeepSeek V4 Pro leads bounded generation (80.6% SWE-bench Verified, 93.5% LiveCodeBench, #1 globally). The open frontier is within a few points of the closed one.

Is it a drop-in for the OpenAI API?

Yes. Change the base URL and key and Cursor, Zed, OpenCode, LangChain or your own clients keep working unchanged.

Can we self-host?

Yes. Dedicated or on-premise runs the same models and API inside your own infrastructure, so customer code never leaves your perimeter.

// get started

START BURNING TOKENS

Skip the AI infra work. Deploy your first private inference endpoint today.

Flat rate. EU data. OpenAI API compatible.