JustVugg/colibri: Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦 - one of the wildest projects recently, because it can run a frontier model on completely unsuitable hardware. With abysmal tokens/second, but it runs. And there are already Metal and CUDA backends, so there's already some acceleration.