A short series on what happens when you run local language models on CPU-only homelab hardware. The first post builds a benchmark suite and finds that most models are either fast or smart, but not both. The second adds context-heavy benchmarks and shows why the speed numbers from a three-sentence prompt lie to you.
I added context-heavy benchmarks to ollama-bench and watched tok/s fall off a cliff. Here's what 8K tokens of system prompt does to 14 models on a dual Xeon.