open model guide · september 2026

The right open model for every use case.

82 enterprise use cases and the open-weight model that solves each one. No closed APIs: everything here can be downloaded, audited and run under your control.

quick selector

Two questions, one model.

Pick the type of task and your main constraint. The recommendation updates instantly; the 82 cases in detail are just below.

What kind of task?

What is your main constraint?

Qwen3.8-27BAPACHE 2.0

The newest of the Qwen generation and the strongest open writer you can run on one machine: 52 on the Artificial Analysis index, Apache 2.0, and native multilingual with solid Spanish.

Apache 2.0262K→1M ctx27B denseimage + video coming soon to Helmcode self-host or provider

see_the_cases_for_this_model →

the 82 cases

Every use case, its open model.

Filter by area or complexity. Each card gives the primary recommendation, the alternative and why. The badges under the model name say where you can run it without hosting it yourself: in Helmcode today, on the NaN.builders cluster, or soon in Helmcode.

11 open models cover the 82 cases. Click one to see only its own.

area
complexity
model:
01
Universal Low

Summarizing long emails

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Low complexity and high volume, so the bill decides: $0.04 per task with 3.8B parameters active, and 256K of context holds a thread nobody wants to read twice.

Open alternative: DeepSeek V4 Flash, for threads that need the 1M window

Use case: Summarization →

02
Universal Medium

Writing professional emails

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

The whole company writes emails all day, and often not in one language, so coverage and cost per message decide: more than 35 languages out of the box, 140 in pretraining, and $0.04 per task with 3.8B parameters active. Multilingual by design, fine-tunable on how your team already writes, and on our flat rate.

Open alternative: Qwen3.8-27B, 27B dense on one machine

Use case: Content generation →

03
Universal Low

Summarizing very long documents

Recommended open model

Qwen3.8-Flash-Next

available on NaN.builders coming soon to Helmcode

Qwen Community 1.0 · 262K→1M ctx · 125B/6B act · image + video

1M of context takes a whole report, case file or book in one pass, with no chunking, and 56 on the Artificial Analysis index is what tells you it understood it. A window only says what fits. Its licence is not a free and open-source one, so read it.

Open alternative: DeepSeek V4 Flash, the same window at the market’s price floor

Use case: Summarization →

04
Universal Medium

Transcribing and summarizing meetings

Recommended open model

Whisper large-v3-turbo

available in Helmcode · Whisper large-v3

MIT · 809M · 99 languages · real time

Whisper large-v3-turbo transcribes the whole recording with no length cap, at 809M parameters and MIT, and V4 Flash turns the transcript into minutes.

Open alternative: Whisper large-v3 on our flat rate, 4.1% word error rate against turbo’s 4.6%

Use case: Voice & transcription →

05
Universal Low

RAG over internal documentation

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

The context is already retrieved: extraction plus writing with precise citations. For confidential docs, self-hosting is the only valid option.

Open alternative: Qwen 3.6 35B

Use case: RAG over internal knowledge →

06
Universal Low

Prioritizing the inbox

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Classification against clear criteria in milliseconds, on one GPU, over every email that comes in. Only 3.8B of its 25.2B parameters fire per token, which is what keeps the bill per message down.

Open alternative: DeepSeek V4 Flash

Use case: Classification & triage →

07
Universal Low

Creating tasks from emails and chats

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Entity extraction (what, who, when) with reliable JSON, over everything the team writes in a day. At that volume cost decides, and 3.8B active parameters are the cheapest tier on the board.

Open alternative: DeepSeek V4 Flash

Use case: Autonomous agents →

08
Universal Low

Automated email follow-ups

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Follow-ups should read as human, not automated. Same 27B dense footprint as the Qwen3.6-27B it replaces, and fourteen points higher on the index than it.

Open alternative: Qwen 3.6 35B, the MoE we serve on a flat rate

Use case: Autonomous agents →

09
Universal Medium

Proofreading and improving text

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Editing means respecting the author’s voice instead of rewriting it. Multilingual by design, 27B dense so it fits one machine, and cheap enough to run over everything the team writes.

Open alternative: GLM-5.3-Flash, or Qwen 3.6 35B on a flat rate

Use case: Content generation →

10
Universal Low

Bulk translation

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

No open model on this page covers more language for less: more than 35 languages out of the box, 140 in pretraining, Apache 2.0, and 3.8B parameters active per token, which is what makes translating a whole catalogue affordable.

Open alternative: Qwen3.8-27B for fine cultural adaptation, or V4 Flash at the highest volumes

Use case: Translation →

11
Universal Low

Semantic search over the KB

Recommended open model

qwen3-embedding + rerank

available in Helmcode

Apache 2.0 · embeddings

The forgotten half of every RAG stack: good embeddings plus reranking decide more than the generator model does.

Open alternative: DeepSeek V4 Flash (synthesis)

Use case: Semantic search & recommendation →

12
Universal Low

Reranking search results

Recommended open model

qwen3-embedding + rerank

available in Helmcode

Apache 2.0 · embeddings

Reordering the retrieved top-k multiplies RAG precision at a minimal marginal cost.

Open alternative: Gemma 4 26B as cross-encoder

Use case: Semantic search & recommendation →

13
Universal Low

Generating FAQs from documentation

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Comprehension plus rephrasing into user language. Ideal for manuals and policies that cannot leave for external APIs.

Open alternative: Qwen3.8-27B

Use case: Content generation →

14
Universal Medium

Executive reports

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Data synthesis plus narrative: the best open reasoning following complex structured templates, and since 26 August the best of them costs a fifth of what it did.

Open alternative: GLM-5.3, which we serve on a flat rate

Use case: Content generation →

15
Universal Medium

Presentations from data

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

One conclusion per slide, title as message. Wired to python-pptx or reveal.js: a full data-to-deck pipeline.

Open alternative: GLM-5.3 on our flat rate, or Qwen 3.6 35B on one machine

Use case: Content generation →

16
Universal Medium

Policies and procedures

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Precise language and consistent terminology across long documents. Internal content that should stay at home.

Open alternative: GLM-5.3-Flash, or Qwen 3.6 35B on a flat rate

Use case: Content generation →

17
Universal Medium

Meeting agent (agenda + minutes)

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

A 1-3h meeting is 50-150K tokens, so 1M of context holds a whole week of them. Whisper large-v3-turbo transcribes and V4 Flash writes the agenda and the minutes, both self-hosted.

Open alternative: GLM-5.3-Flash for meetings that end in a decision

Use case: Autonomous agents →

18
Universal Low

News and intelligence summaries

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

RSS/scraping → batch summary and classification → dashboard. Cost per article is practically zero.

Open alternative: Gemma 4 26B

Use case: Summarization →

19
Universal Low

Detecting urgency in messages

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

One of the simplest tasks on the list: millisecond latency, on-premise, on one GPU, with 3.8B of its 25.2B parameters active per token.

Open alternative: DeepSeek V4 Flash

Use case: Classification & triage →

20
Universal Low

Project pipeline management

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Task status, alerts and summaries. Milestones and owners of client projects never leave the internal network.

Open alternative: Qwen3.8-27B

Use case: Autonomous agents →

21
Universal Medium

Research over your own documents

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

The best open reasoning for synthesizing sources you have already gathered, with 1M of context to hold them. Pair it with your own scraping for the capture step.

Open alternative: GLM-5.3, or DeepSeek V4 Pro 0813

Use case: Research →

22
Code & IT High

Code generation

Recommended open model

GLM-5.3

available in Helmcode

GLM-5.3 licence (custom) · 1M ctx · 744B · text only

Weights public since 25 August, and it matches Kimi K3 on half the weights: 88.2 on Terminal Bench 2.1 against K3’s 88.3, and 28.3 on Terminal Bench 3.0 against its 17.4, Z.ai’s own figures. We serve it on a flat rate, so the two things to plan for only apply if you run it yourself: 755 GB, about two 8-GPU nodes, and a licence that is MIT plus a clause aimed at $10bn model providers.

Open alternative: Kimi K3, or GLM-5.3-Flash when the budget is one node

Use case: Professional copilots →

23
Code & IT High

Autonomous coding agents

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

The open leader in agentic coding, and the lead grew on 26 August: Z.ai publishes 84.3 on Terminal Bench 2.1 and 63.4 on DeepSWE, its own figures, against 81.0 for GLM-5.2.

Open alternative: GLM-5.3, which we serve on a flat rate

Use case: Autonomous agents →

24
Code & IT Medium

Automated code review

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Reasons over the full diff to spot bugs, debt and security issues. With every pull request going through it, price decides: $0.11 per task against V4 Pro’s $0.27, and nothing leaves your infrastructure.

Open alternative: GLM-5.3 on our flat rate when the review is about security (84.5 on CyberGym, the best open figure), or DeepSeek V4 Pro 0813 for the hardest diffs

Use case: Professional copilots →

25
Code & IT Medium

Test generation

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Edge-case coverage and coherent mocks, at the lowest price per task in the open field. DeepSeek publishes that Flash 0731 beats the V4 Pro preview on all nine of its agent benchmarks.

Open alternative: GLM-5.3-Flash, or DeepSeek V4 Pro 0813

Use case: Professional copilots →

26
Code & IT High

Long migrations and refactors

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Hours-long multi-file refactors, where the model holds the repo and its own plan at once: 1M of context, and 328 GB of weights against the 1.5 TB GLM-5.2 needed for the same job.

Open alternative: GLM-5.3, which we serve on a flat rate

Use case: Code migration →

27
Code & IT Low

SQL from natural language

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Text-to-SQL over your schema with a few-shot prompt. This is code, so the coding generation matters: Alibaba publishes 90.3 on LiveCodeBench v6. The schema is sensitive: better kept at home.

Open alternative: Qwen 3.6 35B, or V4 Flash

Use case: Professional copilots →

28
Code & IT Medium

Automated technical documentation

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Reads the code and writes the docs, holding coherence across modules in a large codebase. Z.ai publishes 84.3 on Terminal Bench 2.1 for that long multi-step work, with 1M of context.

Open alternative: GLM-5.3 on a flat rate, or Qwen3.8-27B on one machine

Use case: Content generation →

29
Code & IT Low

L1 support / IT helpdesk agent

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Resolves 80% of L1 queries by loading the whole KB into context. Employee and systems data: inside the network.

Open alternative: Gemma 4 26B

Use case: Customer support →

30
Code & IT High

Log analysis and root cause (RCA)

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Multi-step reasoning connecting unrelated events, with a method (5 Whys, Ishikawa). Confidential logs: self-host.

Open alternative: GLM-5.3, or DeepSeek V4 Pro 0813

Use case: Autonomous agents →

31
Code & IT Medium

Frontend and UI from mockups

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Reads the mockup and writes the component, with native vision and current-generation coding in the same 27B weights. It needed a 64-GPU cluster until August; now it needs one machine.

Open alternative: Qwen3.8-Flash-Next if the mockups justify a cluster

Use case: Professional copilots →

32
Code & IT Low

Synthetic data generation

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Training and test datasets at zero marginal cost. The MIT license places no restriction on how you use the outputs.

Open alternative: Qwen 3.6 35B

Use case: Synthetic data →

33
Code & IT Low

PII detection and anonymization

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Fine-tuned on your own categories of personal data, it runs at the edge of the pipeline before anything leaves.

Open alternative: DeepSeek V4 Flash

Use case: Anonymization →

34
Voice & Multimodal Low

Call transcription (STT)

Recommended open model

Whisper large-v3-turbo

available in Helmcode · Whisper large-v3

MIT · 809M · 99 languages · real time

Pure transcription at call-centre volume, and the model everyone actually runs wins: 809M parameters, MIT, 99 languages and 4.6% word error rate on Artificial Analysis.

Open alternative: Whisper large-v3 when accuracy decides, on our flat rate

Use case: Voice & transcription →

35
Voice & Multimodal Low

Voicebots and speech synthesis (TTS)

Recommended open model

Chatterbox Multilingual

MIT · 500M · 23 languages · voice cloning

23 languages, Spanish among them, voice cloning from a few seconds of reference, MIT and 500M parameters. A full open voicebot: Whisper → LLM → Chatterbox. Kokoro only speaks English.

Open alternative: Kokoro, for English only and the smallest footprint

Use case: Voice & transcription →

36
Voice & Multimodal Low

OCR + extraction from scans

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Reads the document image and returns the fields as JSON in one pass, with up to 16,384 visual tokens per page to keep the small print readable, which is why the pick here is the LLM you already run. If OCR is the whole job, a specialist beats it: GLM-OCR is 0.9B under MIT and leads OmniDocBench v1.5 at 94.62.

Open alternative: GLM-OCR when the pipeline is pure OCR, Qwen3.8-Flash-Next for the highest accuracy tier

Use case: Document extraction →

37
Voice & Multimodal Medium

Product image analysis

Recommended open model

Qwen3.8-Flash-Next

available on NaN.builders coming soon to Helmcode

Qwen Community 1.0 · 262K→1M ctx · 125B/6B act · image + video

Visual QA, catalog tagging and asset verification. Alibaba publishes 88.5 against the 27B’s 85.9 on real-world perception, which is the axis this case turns on, and 6B active out of 125B keeps a whole catalogue affordable.

Open alternative: Qwen3.8-27B on one machine, or Gemma 4 26B on a flat rate

Use case: Visual inspection →

38
Voice & Multimodal Low

Multimodal moderation

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Text and image in the same call, fine-tunable to your platform’s criteria for more consistency than zero-shot.

Open alternative: Qwen3.8-27B, Apache 2.0 and fine-tunable

Use case: Visual inspection →

39
Sales High

Sales proposals from a briefing

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Persuasive writing with frameworks (SPIN, Challenger) and a customer-centered narrative, and enough head to hold the argument across a long proposal.

Open alternative: GLM-5.3-Flash, or the full GLM-5.3 on a flat rate

Use case: Content generation →

40
Sales Low

Lead scoring and prioritization

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Fine-tuned on your conversion history, it beats zero-shot by a wide margin: the structural advantage of open source.

Open alternative: DeepSeek V4 Flash

Use case: Classification & triage →

41
Sales Low

Opportunity tracking

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Briefings from CRM activity. Amounts and clients under negotiation: commercial intelligence that must not leave.

Open alternative: Qwen3.8-27B

Use case: Autonomous agents →

42
Sales Medium

Sales scripts and calls

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Objection handling and personalization by profile. The sales script is confidential strategy: self-host.

Open alternative: Qwen 3.6 35B on a flat rate, or GLM-5.3-Flash for the hardest accounts

Use case: Content generation →

43
Sales High

Negotiation preparation

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

The most solid open strategic reasoning that fits on one machine: 57 on the Artificial Analysis index and 328 GB of weights, against the 60 of GLM-5.3 at 755 GB and of Kimi K3 at 1.56 TB.

Open alternative: GLM-5.3, which we serve on a flat rate

Use case: Research →

44
Marketing Medium

Marketing content

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Creativity, copywriting and adaptation to brand tone. At high volume, the saving over a closed API is substantial.

Open alternative: Qwen 3.6 35B on a flat rate at the highest volumes

Use case: Content generation →

45
Marketing Low

Copy and A/B variations

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Dozens of variations per campaign at zero marginal cost, and it reads the creative as well as the brief: native vision over image and video comes in the same weights.

Open alternative: Qwen 3.6 35B, or V4 Flash at the highest volumes

Use case: Content generation →

46
Marketing Low

SEO: briefs and meta descriptions

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Structured batch generation across the whole catalog or blog. High volume, clear criteria: cost is what rules.

Open alternative: Qwen3.8-27B

Use case: Content generation →

47
Marketing Low

Review sentiment analysis

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Thousands of opinions in batch with aspect-based analysis, at $0.04 per task. Fine-tuning it on your sector’s vocabulary is what sharpens the result, and that is only possible on open weights.

Open alternative: DeepSeek V4 Flash

Use case: Classification & triage →

48
Marketing Low

Content moderation

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Your platform’s own criteria, fine-tuned, at a practically zero cost per message, and text and image go through the same call.

Open alternative: Qwen3.8-27B (fine-tune)

Use case: Classification & triage →

49
Customer support Low

Ticket classification and routing

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Real-time multi-label, fine-tunable on your ticket history for business-specific precision.

Open alternative: DeepSeek V4 Flash

Use case: Classification & triage →

50
Customer support Low

Automated customer replies (FAQ)

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

A support chatbot with customer data under GDPR: in regulated sectors, self-hosting is not optional.

Open alternative: Qwen3.8-27B

Use case: Customer support →

51
Customer support Medium

Handling complaints

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Empathy plus firmness plus a concrete solution. Complaints contain personal data: process it at home.

Open alternative: Qwen 3.6 35B on a flat rate, or GLM-5.3-Flash

Use case: Customer support →

52
Customer support Low

Survey and NPS analysis

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Thematic classification and sentiment over thousands of open responses, at the lowest cost per response on the board, with the personal data never leaving EU infrastructure.

Open alternative: DeepSeek V4 Flash

Use case: Classification & triage →

53
Customer support Low

Support conversation summaries

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Automatic closure of each ticket with a summary for the CRM. High volume, sound writing, one machine.

Open alternative: Qwen 3.6 35B, or V4 Flash at the highest volumes

Use case: Summarization →

54
HR Medium

CV and candidate screening

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Telling evidence from generic claims, without bias. Candidate data is GDPR territory: deploy locally.

Open alternative: GLM-5.3-Flash, or the full GLM-5.3 on a flat rate

Use case: Document review & QA →

55
HR Low

New employee onboarding

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

All the onboarding documentation fits in 1M of context, with no RAG and without internal policies leaving.

Open alternative: Gemma 4 26B

Use case: RAG over internal knowledge →

56
HR High

Performance reviews

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Constructive feedback with nuance. This data can end up in labor proceedings: self-host plus encryption.

Open alternative: GLM-5.3-Flash, or the full GLM-5.3 on a flat rate

Use case: Content generation →

57
HR Low

HR chatbot for employees

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

HR data is the most sensitive in the company: the conversations do not leave the corporate infrastructure.

Open alternative: Qwen3.8-27B

Use case: Professional copilots →

58
HR Low

Job postings

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Persuasive writing with a defined structure, in native Spanish, even for confidential roles.

Open alternative: Qwen 3.6 35B

Use case: Content generation →

59
HR Medium

Training and educational materials

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Instructional design adapted to the learner’s level. In high-volume e-learning, cost per module tends to zero.

Open alternative: Qwen 3.6 35B, the MoE we serve on a flat rate

Use case: Content generation →

60
HR Low

Internal comms and newsletters

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Good tone and structure on one machine. Internal communications stay out of third-party APIs.

Open alternative: Gemma 4 26B, or Qwen 3.6 35B

Use case: Content generation →

61
Finance Low

Invoices and expenses (extraction)

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Processes the invoice image directly and returns structured JSON, with up to 16,384 visual tokens per page against the 280 of the Gemma we serve. A scanned A4 is where the small print decides the amount. A 100% internal pipeline on one machine, and if that pipeline only ever reads documents, a specialist such as GLM-OCR (0.9B, MIT) is more efficient still.

Open alternative: Gemma 4 26B, which we serve on a flat rate

Use case: Document extraction →

62
Finance High

Financial analysis and reporting

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

One index point behind V4 Pro at less than half the cost, with the same 1M of context for a full set of accounts and its notes. For listed companies or M&A, self-hosting is close to mandatory.

Open alternative: DeepSeek V4 Pro 0813, or GLM-5.3-Flash

Use case: Document analysis →

63
Finance High

Risk and fraud detection

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Complex patterns over the full transaction history, and the 1M of context is the model’s own published number. In banking the data stays in, so MIT downloadable weights are what count.

Open alternative: GLM-5.3, or DeepSeek V4 Pro 0813

Use case: Classification & triage →

64
Finance Low

Bank reconciliation and matching

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

High volume of short comparisons: statement lines against ledger entries, with the near-matches flagged for a person. Cost per token is what decides here.

Open alternative: Gemma 4 26B

Use case: Document review & QA →

65
Finance High

Credit analysis and underwriting

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Reads the accounts and writes the reasoning behind a limit, ratio by ratio, at $0.11 per file. A solvency judgement on a named company belongs self-hosted, and a person still signs it.

Open alternative: GLM-5.3-Flash for the files that go to committee

Use case: Document analysis →

66
Finance Medium

Collections and payment reminders

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

The same message has to escalate from a nudge to a formal notice without losing the client, so register weighs more than raw reasoning. 27B dense weights under Apache 2.0.

Open alternative: Qwen 3.6 35B, the MoE we serve on a flat rate

Use case: Content generation →

67
Finance Low

Expense policy checks

Recommended open model

Gemma 4 26B

available in Helmcode

Apache 2.0 · 256K ctx · 25.2B/3.8B act · multimodal

Judgement against clear criteria, which is what this size is for: this dinner is over the per diem, this taxi has no attendees. The policy travels in the prompt and changes without retraining anything, at $0.04 per task.

Open alternative: DeepSeek V4 Flash

Use case: Document review & QA →

68
Legal High

Contract and clause analysis

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

The highest open level of nuance comprehension you can deploy, with MIT license and model card published: auditable end to end.

Open alternative: GLM-5.3, or DeepSeek V4 Pro 0813

Use case: Document review & QA →

69
Legal High

Compliance and regulation

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

With the AI Act applying in general since 2 August 2026, an open, traceable stack on EU infrastructure gives your assessment the evidence it needs.

Open alternative: GLM-5.3, which we serve on a flat rate

Use case: Document review & QA →

70
Legal Low

Legal document classification

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Tagging and triage of case files by matter, jurisdiction and urgency, without a single page leaving the firm.

Open alternative: Qwen 3.6 35B on a flat rate, or Gemma 4 26B

Use case: Classification & triage →

71
Legal High

Due diligence in a data room

Recommended open model

Kimi K3

K3 licence (custom) · 1M ctx · 2.8T/104B act · multimodal

Thousands of documents that only mean something read together, so reasoning across them matters more than window size: K3 leads every open model on AA-LCR at 82.7%.

Open alternative: GLM-5.3 on our flat rate, the same index score on half the weights, or DeepSeek V4 Flash on hardware you already have

Use case: Document analysis →

72
Legal High

Case law and precedent research

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Reasoning that has to hold a chain of citations without inventing one. Ground it with RAG on your own database of rulings, never on the model memory.

Open alternative: GLM-5.3, or DeepSeek V4 Pro 0813

Use case: Research →

73
Legal Medium

Drafting from a clause library

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Generation rather than analysis: assembling a first draft out of clauses legal has already approved, so the team edits instead of starting from nothing.

Open alternative: GLM-5.3-Flash, or Qwen 3.6 35B on a flat rate

Use case: Content generation →

74
Legal Medium

Regulatory change monitoring

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Watching what changed in a regulation and which internal policies it touches. A different job from checking compliance: this one runs before anybody is out of it.

Open alternative: DeepSeek V4 Flash

Use case: Autonomous agents →

75
Operations & Strategy Low

Supplier and procurement management

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Comparing terms and flagging expirations. Agreed prices are confidential commercial information.

Open alternative: GLM-5.3-Flash (every contract at once, 1M ctx)

Use case: Autonomous agents →

76
Operations & Strategy Medium

Competitor and market analysis

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

Structured reasoning over sources you have already gathered. For web capture, combine it with your own scraping.

Open alternative: DeepSeek V4 Flash

Use case: Research →

77
Operations & Strategy Medium

RFP and tender responses

Recommended open model

Qwen3.8-27B

coming soon to Helmcode

Apache 2.0 · 262K→1M ctx · 27B dense · image + video

Hundreds of requirements answered one by one, each traceable to a document you already hold. The bid stays confidential until the envelope is opened.

Open alternative: GLM-5.3-Flash, or the full GLM-5.3 on a flat rate

Use case: Content generation →

78
Operations & Strategy Medium

Security and vendor risk questionnaires

Recommended open model

GLM-5.3-Flash

available on NaN.builders coming soon to Helmcode

MIT · 1M ctx · 320B/18B act · image + video

The same two hundred questions arrive from every client with the wording changed. Answered from your own evidence base, with a person signing it off.

Open alternative: DeepSeek V4 Flash

Use case: Document review & QA →

79
Operations & Strategy Low

Logistics exception triage

Recommended open model

DeepSeek V4 Flash 0731

available in Helmcode

MIT · 1M ctx · 284B/13B act

Thousands of short events a day, each needing a route: delay, damage, wrong address or nothing at all. Volume picks the model.

Open alternative: Gemma 4 26B

Use case: Classification & triage →

80
Operations & Strategy Medium

Quality inspection from photos

Recommended open model

Qwen3.8-Flash-Next

available on NaN.builders coming soon to Helmcode

Qwen Community 1.0 · 262K→1M ctx · 125B/6B act · image + video

Reads the picture off the line, or the clip, and writes the defect report against your own criteria. Alibaba puts it at 88.5 on real-world perception, its best published figure of any model it compares itself against, and a missed defect costs more than the inference does.

Open alternative: Qwen3.8-27B on one machine, or Gemma 4 26B on a flat rate

Use case: Visual inspection →

81
Marketing Low

Campaign and social imagery

Recommended open model

FLUX.2 [klein] 9B

available in Helmcode available on NaN.builders

FLUX Non-Commercial · 9B · ~29 GB VRAM · image out

Dozens of pieces a week with no stock licence and no agency round: 9B parameters, about 29 GB of VRAM, and a generation in under a second. On our flat rate, and on the NaN.builders cluster. Its licence is not a free and open-source one, so read it before it goes into a commercial pipeline.

Open alternative: FLUX.2 [klein] 4B, the Apache 2.0 route on one consumer GPU, or Qwen-Image-2512

Use case: Content generation →

82
Marketing Medium

Product imagery and creative variants

Recommended open model

FLUX.2 [klein] 9B

available in Helmcode available on NaN.builders

FLUX Non-Commercial · 9B · ~29 GB VRAM · image out

Backgrounds, formats and versions of the same shot: editing takes several reference images and runs on the same weights that generate, so a catalogue is one model and not two pipelines. On our flat rate.

Open alternative: Qwen-Image-2512, Apache 2.0, when the piece leans on text inside the image

Use case: Content generation →

available in Helmcode (flat rate, EU, zero logs) available today on the NaN.builders cluster announced for Helmcode, not serving yet complexity = how demanding the task is, not the deployment Kimi K3, GLM-5.3, Qwen3.8-Flash-Next and FLUX.2 [klein] 9B ship under their own terms, with use restrictions: check them before you build

The gap with closed models only exists at the frontier edge. Almost no real work lives there.

In production 99.5% of our tokens go through open models, and not on the easy tasks: classifying, extracting, summarizing, drafting, answering and reasoning over documents nobody else has read. What is left is a narrow set of genuinely frontier problems, and the answer to those is to route them to the big model, not to pay frontier prices for the other 99.5%.

get started

Start with one case, not with a migration.

Pick a case from the list, point the same code at our endpoint and read both outputs side by side.

The API you already call. One afternoon. Nothing to migrate.

A process with volume and stable criteria can become a small, specialized model with the weights in your name, running on far less hardware. model_specialization →