Well That's Just Plain Silly

So FreeToken makes local-first options look insane now. not only am I seeing deepseek flash being run on two 3090s local, but also GPUs like the 3080 no longer appear to be priced appropriately!

Imagine: a 3080 running MoE models hosted mostly in RAM, you can and would be able to carry long-context operations with speculative decoding at rather-high TPS simply by allowing the used expert to reside on VRAM. Whole worlds of Local Language Models open up at silly prices because we are no longer tethered to jamming compressed models directly onto the GPU.

I, for one, am looking at the used market a bit more closely now and am reconsidering what self hosting has to look like. And, if large enterprise ram busses are on the table, there is a possibility we could be looking at higher TPS for multi-agent long-context boxes running with compartively-tiny VRAM counts for what was only a week ago looking more and more bleak in terms of self hosting.

Watershed moment, at least until the 3080s get bought up, my link leads to an empty seller page, and we are back where we used to be in a GPU drought. Hey, at least it's something. Maybe RAM prices will go down if VRAM stops hogging all of the manufacturing attention!

I won't hold my breath!

← All notes