free tool · no signup · runs in your browser

will it run?

pick your machine. see which open ai models fit, and roughly how fast they'll talk. before you download 20GB and find out the hard way.

your machine

how it works

two numbers decide it: memory and bandwidth.

fit. the model file, plus working memory for the conversation, plus about 1GB for the runtime, has to fit in the memory your gpu can use. on a mac that's roughly two thirds of your unified memory by default.

speed. to write each word, the model reads its active weights from memory once. so speed is capped by memory bandwidth divided by model size. we checked this on a real M3 Pro 18GB: Qwen 3.5 9B Q4 ran at 19.8 tokens a second, inside the range this page predicts.

model sizes come from the ollama library. bandwidth comes from apple, nvidia and amd's published specs.

questions

why does a mac only get about two thirds of its memory?

macos lets the gpu wire roughly two thirds of unified memory on machines with 36GB or less, and about three quarters above that. the rest stays with the system and your apps. you can raise the limit, but the defaults are what most people run, so that's what we count.

what do q4 and q8 mean?

how many bits each weight is stored in. q4 is about a quarter the size of the original model and the usual choice at home: small quality loss, big savings. q8 is roughly twice the size of q4 and a little closer to the original. if q4 fits comfortably and q8 is tight, q4 is almost always the better trade.

why are the moe models so fast for their size?

a mixture-of-experts model like qwen 3.6 35b-a3b stores 35 billion parameters but only uses about 3 billion for each token it writes. it needs the memory of a big model and runs at close to the speed of a small one.

what does "tight" mean?

it fits, but with little room left over. it will run if you close other heavy apps, and a longer conversation (more context) can push it over. comfortable means there's real headroom.

how accurate are the speeds?

they're estimates from memory bandwidth, which is what limits how fast a local model writes. we checked the method on a real machine: an M3 Pro 18GB running Qwen 3.5 9B Q4 measured 19.8 tokens a second, inside our estimated range. your numbers depend on the runtime, context length and what else is running.

which runtime should i use?

we set people up on llama.cpp with the model as a plain gguf file and open webui as the chat app. ollama works too and is the quickest way to try a model, which is why the commands below use it.