Aller au contenu principal
Nicolas Cousin Tech SolutionsNicolas Cousin Tech Solutions

Local LLMs on CPU: the Cold Shower

Numbers, not a feeling: what a local LLM actually delivers on a CPU-only VPS, compared to the cloud. Fourth article in the series.

By Nicolas Cousin — Published on October 6, 2026

Local LLMs on CPU: the Cold Shower

TL;DR

On the series' VPS (8 vCores, 24GB RAM, CPU-only), I timed what Ollama actually delivers, from a tiny prompt to a full coding task, and compared it to the same type of question asked of a cloud model. Result: a question that takes 3 seconds in the cloud took 4 min 44 locally via Hermes, and a coding task went past a quarter of an hour. The VPS does run 7B to 24B models, but "it runs" doesn't mean "it's interactive."


Table of contents


What I measured, and what I didn't

This isn't a scientific model-versus-model benchmark: the hardware, the cloud provider, and the pipeline differ from one row of the table below to the next. These are times observed under real usage conditions, on my CPU-only VPS, while I was building the environment described in the previous articles. As a measure of user experience, meaning what it actually feels like to wait for a reply, it's telling:

Real test Model Observed time Note
Tiny direct prompt (~18 input tokens) Mistral NeMo ~36 s Already noticeable for a trivial interaction
Direct prompt, ~2,225 input tokens Mistral NeMo ~2 min 20 s Prefill cost becomes obvious
Simple .NET question via Hermes Mistral NeMo ~4 min 44 s Orchestration + Hermes context worsen latency
Substantial coding reply Devstral 24B ~16 min 27 s First real cold shower
Multi-file task Devstral 24B ~18 min 28 s Result wasn't usable as real agent work
Equivalent simple question GPT-5.6 Sol (cloud) ~3 s Radically different order of magnitude

The cost of prefill, visible even on small prompts

The first row of the table alone doesn't say much: 36 seconds for ~18 input tokens is slow, but you could assume it's just Mistral NeMo generating slowly. The second row changes the reading: bumping the input to ~2,225 tokens, without changing anything else, pushes the time up to ~2 min 20.

These two measurements don't explicitly separate the time spent ingesting the prompt (prefill) from the time spent generating the reply: in theory part of the gap could come from a longer reply on the second test. But the order of magnitude strongly suggests prefill carries real weight: before producing a single token of the reply, the model first has to ingest the whole prompt, and on CPU that step is expensive, one that grows with the size of the input context, not just the size of the expected reply.

Hermes adds its own layer of latency

The third row points the same direction as the previous article: a comparable question asked to Mistral NeMo, but through Hermes rather than calling Ollama directly, takes ~4 min 44, more than the raw prompt at a comparable context size from the previous row (~2 min 20). This isn't a controlled measurement (the two prompts aren't identical), but the gap goes the way you'd expect: Hermes injects its own system prompt, tool definitions, and conversation history before even getting to the question itself, and on CPU, each of these layers costs extra seconds of prefill.

To put a scale on that number: the same question asked of a cloud model (GPT-5.6 Sol) takes about 3 seconds. Roughly 95 times longer in this specific test, for a basic interactive use case.

Devstral 24B: the real cold shower

The previous times were for simple questions. On an actual coding task, with Devstral 24B, the time climbs to ~16 min 27 s for a substantial reply, and ~18 min 28 s for a task spread across multiple files, without the result even being usable as autonomous agent work, a topic I'll come back to in more detail in an article dedicated to those agentic benchmarks.

This is where the question changes nature. Under a few minutes, you can still call it a usability discomfort. Past a quarter of an hour for a task that would take a few seconds to a few minutes in the cloud, it stops being an ergonomics problem: it's an entirely different mode of use you have to consider.

What memory tells you

CPU isn't the only factor to watch on a 24GB-RAM VPS. A few orders of magnitude for the memory footprint of the models handled at this stage of the series:

  • Devstral 24B: ~14GB
  • GPT-OSS 20B: ~13GB
  • Qwen3-Coder 30B: ~18GB
  • GLM-4.7-Flash: ~19GB

These are model-weight footprints, not the KV-cache (context) cost discussed in the previous article with num_ctx: the two add up on the same 24GB. Models this size fit within the VPS's memory budget, but leave proportionally less room for context the more of that budget their weights already occupy.

What this changes

A CPU VPS at under €30 a month can genuinely run 7B to 24B LLMs: that's not an empty promise, the models load, run, and reply. But "it runs" doesn't mean "it's interactive." A question that took a few seconds in the cloud took several minutes locally, and a coding task went past a quarter of an hour.

This result doesn't invalidate local: what it invalidates is the original idea of an interactive copilot running entirely on this CPU. That distinction, between interactive use and deferred use, is what structures what comes next: what Hermes does with its background mode, once you stop demanding an immediate reply.