AI PC vs GPU vs Mac: what actually runs a local LLM?
Three hardware paths, one hype-free answer — the NPU your "AI PC" advertises is not the part that runs a local chatbot. Here is what does.
Head-to-head on the things you'd actually decide between.
8 articles
Three hardware paths, one hype-free answer — the NPU your "AI PC" advertises is not the part that runs a local chatbot. Here is what does.
They are the same engine. You are choosing an interface and a licence, not a speed.
Output costs five to six times more than input — at every vendor, at every tier. That ratio, not the headline price, is what decides your bill.
Google gives away every Flash text model and an older Pro. It does not give away the current Pro, or a single image.
Image APIs bill in tokens, not images. Only one major vendor will tell you what a picture costs.
Audio tokens cost eight times what text tokens cost. But the cache discount on them is close to 99%.
Four current models, one 1M-token window that is not the same size on all of them, and a tokenizer change nobody announced as a price rise.
Five tools, five different ways of quoting a price — and one of them shows different numbers to different people.