What is an NPU?
The chip your laptop and phone now advertise — what it's for, and why its headline "TOPS" number tells you less than it looks.
Author
Founder & Editor
Fiqhro Dedhen founded and edits NeuralGist. A full-stack developer with five years spent building the machinery businesses actually run on — APIs, integrations, internal tools — Dedhen writes about AI from the position of someone who has to make it work in production, where a good demo counts for nothing. The test applied to every tool covered here is not whether it is impressive, but whether it changes the job.
Writes aboutAI tooling, model releases, consumer technology, software development
The chip your laptop and phone now advertise — what it's for, and why its headline "TOPS" number tells you less than it looks.
Teaching a model how to behave — not what is true.
Asking for the working, and getting a better answer as a side effect.
It is slow, and it also silently gave your model a 4K context window. Both are the same bug.
It is not the model. Ollama picks your context length from your VRAM at startup, and the bottom tier is very small.
The file format behind every local model you have downloaded — and how to read its name.
A model that can take actions in a loop — and the word the industry has worn out.
The current frontier and open-weight language models, with every spec read from a primary source.
Turning meaning into coordinates — the trick that makes semantic search work.
The 2017 architecture underneath essentially every model you have heard of.
When a model states something false with total confidence — and why that is the normal case, not a glitch.
Paying once for the part of your prompt that never changes — and the write fee nobody mentions.
Giving the model the documents instead of hoping it memorised them.
The model's working memory — and the reason a long chat starts forgetting things.
Four levers, ranked by how much they actually save. Switching vendor is not one of them.
Shrinking a model so it fits on hardware you own — and what you give up.
They are the same engine. You are choosing an interface and a licence, not a speed.
Output costs five to six times more than input — at every vendor, at every tier. That ratio, not the headline price, is what decides your bill.
Google gives away every Flash text model and an older Pro. It does not give away the current Pro, or a single image.
Image APIs bill in tokens, not images. Only one major vendor will tell you what a picture costs.
The unit models read, count and bill in — and it is not a word.
Audio tokens cost eight times what text tokens cost. But the cache discount on them is close to 99%.
Four current models, one 1M-token window that is not the same size on all of them, and a tokenizer change nobody announced as a price rise.
Five tools, five different ways of quoting a price — and one of them shows different numbers to different people.