

the 6B actives should easily fit into an 8GB card
That’s not how MoE models work. There are many “expert’ models and there is a static router model (dense) which determines which “expert” models to route the tokens through. What you want at a minimum is the dense portion of the model to be on VRAM and all of the weights to be in RAM/VRAM for best performance.

















So this was the Ox-Alpha model on OpenRouter. That’s cool.