GLM 5.2 orquesta, DeepSeek V4 Flash ejecuta
Cómo diseñar sistemas multi-modelo donde un LLM abierto planifica, otro ejecuta y el contexto que se pasan entre ellos es eficiente.
8 entradas
Cómo diseñar sistemas multi-modelo donde un LLM abierto planifica, otro ejecuta y el contexto que se pasan entre ellos es eficiente.
GLM 5.2, el modelo abierto MIT de 744B, ya corre en la infraestructura europea de Helmcode con zero logs y tarifa plana de 150 euros al mes.
La versión tranquila del AI Act: qué obligaciones te aplican de verdad desde el 2 de agosto, cuáles movió el Ómnibus y qué conviene tener listo.
FP8, NVFP4, H200, B200: qué acelera en hardware cada generación de NVIDIA, cuánta VRAM necesita tu modelo y qué GPU comprar. Con datos de producción reales.
117 billion tokens, 3.68 million requests, 21 countries, and 99.98% uptime. NaN is a community of builders with its own inference infrastructure and a private platform to deploy apps and agents.
In this post I'll walk you through how the community's inference servers are set up: the hardware we use, the stack we run, and the models we serve.
I've spent several hours over several days documenting and optimizing my entire local environment so I can "mechanize" the work I do every day managing infrastructure for multiple startups.
This post isn't meant to be a guide on how to use Clawd, but rather a look at how we're rolling it out at Helmcode to have an AI Agent that helps us with our day-to-day work managing the Cloud infrastructure of multiple startups.
// cookies
Solo usamos cookies estrictamente necesarias para que el sitio funcione. Nada de analítica ni publicidad, nunca — consulta la Política de cookies.
// preferencias