How to fix Ollama not using your GPU
It is slow, and it also silently gave your model a 4K context window. Both are the same bug.
Specific problems, specific versions, actual fixes.
3 articles
It is slow, and it also silently gave your model a 4K context window. Both are the same bug.
It is not the model. Ollama picks your context length from your VRAM at startup, and the bottom tier is very small.
Four levers, ranked by how much they actually save. Switching vendor is not one of them.