free tool · no signup · runs in your browser
will it run?
pick your machine. see which open ai models fit, and roughly how fast they'll talk. before you download 20GB and find out the hard way.
| model | verdict | needs | speed | download | try it |
|---|
speeds are estimates in tokens a second (a token is about three quarters of a word). reading speed is around 5 to 10. anything above 15 feels quick.
get this list by email
we'll send your results, and one more email when a better model fits your machine. nothing else.
giving that model an agent? we tested whether a small local model can stop a hijacked agent from sending email, money or crypto. a 9b on a laptop caught 30 of 30 attacks.
see gatebench →how it works
two numbers decide it: memory and bandwidth.
fit. the model file, plus working memory for the conversation, plus about 1GB for the runtime, has to fit in the memory your gpu can use. on a mac that's roughly two thirds of your unified memory by default.
speed. to write each word, the model reads its active weights from memory once. so speed is capped by memory bandwidth divided by model size. we checked this on a real M3 Pro 18GB: Qwen 3.5 9B Q4 ran at 19.8 tokens a second, inside the range this page predicts.
model sizes come from the ollama library. bandwidth comes from apple, nvidia and amd's published specs.
questions
why does a mac only get about two thirds of its memory?
macos lets the gpu wire roughly two thirds of unified memory on machines with 36GB or less, and about three quarters above that. the rest stays with the system and your apps. you can raise the limit, but the defaults are what most people run, so that's what we count.
what do q4 and q8 mean?
how many bits each weight is stored in. q4 is about a quarter the size of the original model and the usual choice at home: small quality loss, big savings. q8 is roughly twice the size of q4 and a little closer to the original. if q4 fits comfortably and q8 is tight, q4 is almost always the better trade.
why are the moe models so fast for their size?
a mixture-of-experts model like qwen 3.6 35b-a3b stores 35 billion parameters but only uses about 3 billion for each token it writes. it needs the memory of a big model and runs at close to the speed of a small one.
what does "tight" mean?
it fits, but with little room left over. it will run if you close other heavy apps, and a longer conversation (more context) can push it over. comfortable means there's real headroom.
how accurate are the speeds?
they're estimates from memory bandwidth, which is what limits how fast a local model writes. we checked the method on a real machine: an M3 Pro 18GB running Qwen 3.5 9B Q4 measured 19.8 tokens a second, inside our estimated range. your numbers depend on the runtime, context length and what else is running.
which runtime should i use?
we set people up on llama.cpp with the model as a plain gguf file and open webui as the chat app. ollama works too and is the quickest way to try a model, which is why the commands below use it.