GLM 5.2 orchestrates, DeepSeek V4 Flash executes
How to design multi-model systems where one open LLM plans, another executes, and the context they hand each other stays efficient.
8 posts
How to design multi-model systems where one open LLM plans, another executes, and the context they hand each other stays efficient.
GLM 5.2, the 744B MIT open model, now runs on Helmcode's European infrastructure with zero logs and a flat rate of 150 euros a month.
This week you will read a lot of headlines with 35 million euros in big type. It is real, it is the top penalty band of the AI Act, and it is also the worst possible way to understand what happens from 2 August.
FP8, NVFP4, H200, B200: what each NVIDIA generation accelerates in hardware, how much VRAM your model needs, and which GPU to buy. With real production data.
117 billion tokens, 3.68 million requests, 21 countries, and 99.98% uptime. NaN is a community of builders with its own inference infrastructure and a private platform to deploy apps and agents.
In this post I'll walk you through how the community's inference servers are set up: the hardware we use, the stack we run, and the models we serve.
I've spent several hours over several days documenting and optimizing my entire local environment so I can "mechanize" the work I do every day managing infrastructure for multiple startups.
This post isn't meant to be a guide on how to use Clawd, but rather a look at how we're rolling it out at Helmcode to have an AI Agent that helps us with our day-to-day work managing the Cloud infrastructure of multiple startups.
// cookies
We only use strictly necessary cookies to run the site. No analytics, no advertising, ever — see our Cookie Policy.
// preferences