GLM 5.2 orchestrates, DeepSeek V4 Flash executes
How to design multi-model systems where one open LLM plans, another executes, and the context they hand each other stays efficient.
3 posts
How to design multi-model systems where one open LLM plans, another executes, and the context they hand each other stays efficient.
FP8, NVFP4, H200, B200: what each NVIDIA generation accelerates in hardware, how much VRAM your model needs, and which GPU to buy. With real production data.
Before we start, a bit of context. The infrastructure is hosted on AWS and the architecture was based on Serverless services:
// cookies
We only use strictly necessary cookies to run the site. No analytics, no advertising, ever — see our Cookie Policy.
// preferences