AI-NAV · Article
Testing Local LLMs on Consumer GPUs: Quality and Speed

Running an open-weight model on a consumer GPU comes down to two numbers: whether the weights fit in VRAM, and how many tokens per second you are willing to wait for. A 24GB card can usually host a 14B to 32B model at 4-bit quantization, while an 8GB card is realistically limited to the 7B to 8B range. Quality is not automatically worse because it is local, but at the same parameter count a 4-bit build tends to slip on long reasoning chains and multi-step tasks compared with a full-precision hosted version.
Three different things people call local deployment
The phrase "I run it locally" covers three setups with very different limits, and comparing them without separating them produces meaningless results.
The first is true local inference. Weights are downloaded to your machine and the forward pass runs on your own GPU. It works offline and no prompt leaves the machine. You pay for it in VRAM, electricity, and a hard ceiling set by your card.
The second is a local front end pointed at a remote API. The chat window runs on your computer, but the request travels to a server. It feels like a hosted chatbot, and it is not private deployment in any strict sense.
The third is local inference exposed as a service on your own network, so other programs or teammates can call it. This is what teams usually mean by self-hosted deployment.
One test separates them: unplug the network cable. If the model still answers, you are in the first or third category. If it stops, you were only running a client.
Which open models fit which VRAM tier
The question "which models can I run locally" is really a question about your VRAM number. Open weights being downloadable does not mean they will load on your card; quantization sits in between.
Quantization compresses weights from higher to lower precision. Common tiers include 8-bit and 4-bit. A 4-bit build typically cuts memory use to roughly a quarter of the original size, at the cost of precision that shows up first in arithmetic and long reasoning.
As a working range: 8GB of VRAM suits 7B to 8B models at 4-bit; 12GB to 16GB handles the 14B class; 24GB can take a 32B model at 4-bit or a 14B model at higher precision. Pushing into the 70B class generally means multiple GPUs, or a large system memory pool with CPU offload, and speed drops noticeably.
Model releases move quickly, so check the official model card and your inference framework's documentation for the actual memory requirement of a specific checkpoint rather than copying someone else's build list.
What actually determines local generation speed
There is no single answer to "how fast is a local LLM," because three variables interact, and changing any one of them changes the experience substantially.
Parameter count is the first. More parameters means more weights must be read to produce each token, so generation slows down. This is why a 7B and a 32B model differ by multiples rather than percentages.
Quantization level is the second. Lower precision saves memory and reduces memory bandwidth pressure, which usually raises throughput. Push it too far and quality falls, so you are trading accuracy for speed.
Memory bandwidth and offloading are the third. If part of the model is pushed into system memory and computed on the CPU, throughput collapses. You can spot this by watching for offload messages in the inference log, or by checking whether VRAM sits near its limit during generation.
A practical self-test: fix a prompt of roughly 200 words, ask for an answer of roughly 300 words, and time from pressing send to the last token appearing. Divide generated words by elapsed seconds. That number reflects your real experience better than any published score.
Open-source software for individual deployment
The search for easy personal deployment usually means: I do not want to hand-build an environment, I want something I can launch and use. The open-source ecosystem splits into two families.
The first bundles model download, quantized loading, and a chat interface into one desktop application. You pick a model and a quantization level, and the tool fetches and loads it. The trade-off is limited control over advanced parameters, and error messages when VRAM runs short are often vague.
The second is an inference server framework that exposes the model behind a common API shape so other programs can call it. It suits people wiring a model into their own project or sharing one with a team, at the cost of more configuration.
If you are still deciding where a local model fits in your workflow, browsing the AI Chatbots category is a reasonable starting point, and the walkthrough in How to Build AI Agents Without Code: From Templates to Production Deployment covers model connection and orchestration that applies equally to a local service.
How to evaluate answer quality without fooling yourself
The most common mistake is asking a few general knowledge questions, getting correct answers, and concluding quality is fine. General knowledge is exactly where quantization damage hides.
Test in layers instead. Layer one is factual recall, checking for obvious hallucination. Layer two is long-document summarization, checking whether the model extracts the point or just paraphrases. Layer three is multi-step reasoning, such as a conditional arithmetic problem, where quantization gaps surface fastest. Layer four is format adherence, for example requiring a fixed JSON structure and checking for missing fields.
Control your variables: same prompt, same temperature, same maximum output length, changing only the model or the quantization level. Otherwise the difference you measure may come from settings rather than the model.
Also separate "the model is weak" from "the deployment is misconfigured." A context window set too small or a mismatched prompt template will make a capable model look bad. Specific features, pricing, and availability should be confirmed on the official page, since they change over time.
Comparing model sizes and their trade-offs
The table below lines up model scale against typical quantization, so you can shortlist candidates before downloading anything. Memory and speed figures are working ranges; actual results depend on the inference framework, driver version, and GPU architecture.
| Model | Typical VRAM | Good for | Main limitation |
|---|---|---|---|
| Llama 3.1 8B | About 6GB at 4-bit | Everyday Q&A, rewriting | Weak multi-step reasoning |
| Qwen2.5 7B | About 6GB at 4-bit | Chinese Q&A, summaries | Loses detail in long text |
| Gemma 2 9B | About 8GB at 4-bit | General conversation | Uneven Chinese fluency |
| Mistral Small 22B | About 14GB at 4-bit | Writing, coding help | Offloads when VRAM is tight |
| Qwen2.5 32B | About 20GB at 4-bit | Complex reasoning, long docs | Noticeably slower |
| DeepSeek-V3 | Multi-GPU or large RAM | Hard reasoning tasks | Not viable on one card |
The pattern: each VRAM tier roughly doubles the model size you can host, but speed does not fall linearly. It breaks sharply once offloading begins. Choose the largest model that still runs without offloading before you chase parameter count.
Prerequisites
Before you start, confirm you have a GPU with a recent driver, at least 8GB of VRAM for a first run, and enough free disk space for a quantized model file, which is typically several gigabytes. You also need a terminal, and on Linux or Windows with WSL, the nvidia-smi utility available on your PATH. If you plan to serve the model to other machines, decide in advance which port it will listen on.
Step-by-step
- Check your GPU and available VRAM. Run
nvidia-smiin a terminal and read the memory total and driver version in the output. If the command is not found, install or repair the driver first; nothing later will work without it. - Install an open-source inference tool. Pick either a desktop runner or a server framework, and follow the install instructions in its official repository. Success looks like a model management screen opening, or a service port reported as listening.
- Download a 7B to 8B model at 4-bit. Search the model library inside the tool for the model name and select the build labeled 4-bit. When the download finishes, the model appears in your local model list.
- Load the model and watch VRAM. Run
nvidia-smiagain, or open your system task manager, and confirm memory use is not sitting at the ceiling. If it is, drop to a smaller model or a more aggressive quantization. - Send a fixed test prompt. Use a prompt of roughly 200 words and set the maximum output length to about 300 words. Record the total elapsed time.
- Convert to a rate. Divide generated words by elapsed seconds to get words per second. This is your baseline for every later comparison.
- Repeat steps 3 to 6 with a larger model. Compare both speed and answer quality to decide which tier your card should settle on.
- Write the result down. Record the model name, quantization level, VRAM used, words per second, and which task types failed. You will reuse this when you switch tools.
How to verify
Verification here means confirming the run was genuinely local and genuinely unoffloaded. Disconnect the network and send another prompt; if the model still answers, inference is happening on your machine. Then check the inference log for offload messages during generation, and confirm VRAM usage stays below the card's limit throughout. Finally, rerun the same prompt twice with identical settings and confirm the timing is stable within a reasonable margin; wild swings usually indicate background processes competing for the GPU.
Troubleshooting
- The model fails to load with an out-of-memory error. The quantization level is too high or the context length is set too large. Switch to a lower-bit build, or reduce the context length and retry.
- Generation suddenly becomes extremely slow. Part of the model is likely being computed on the CPU. Check the log for offload notices, then move to a smaller model or lower precision.
- Output repeats and never stops. Sampling parameters are probably misconfigured, for example a very low temperature with no repetition penalty. Adjust the sampling settings or switch to the prompt template the model card recommends.
- Answers drift into another language or miss the question. The model may have limited training data in your target language, or the prompt template does not match. Try a model with stronger coverage and confirm the template format.
- Other programs cannot reach the service. The server is often bound only to localhost. Check the listen address in the startup arguments, change it to accept LAN connections, and restart.
Frequently Asked Questions
How large a model can a consumer GPU run?
A 24GB card typically handles the 32B class at 4-bit quantization, 16GB suits the 14B class, and 8GB realistically stops at 7B to 8B. Reaching the 70B class usually requires multiple GPUs or CPU offload with a large memory pool, and speed drops sharply. Confirm exact requirements on the official model card.
Is a locally hosted model much worse than a cloud one?
At the same parameter count, a local 4-bit build is close on general knowledge but slips on multi-step reasoning, long arithmetic chains, and strict structured output. The gap comes from quantization loss and context limits, not from being local. Layered testing shows which layer suffers.
How fast does a local LLM need to be to feel usable?
It depends on the task. Reading assistance and rewriting are fine at a modest rate, conversational Q&A feels smooth once output keeps pace with reading speed, and code completion demands much lower latency. Measure your own baseline with a fixed prompt before deciding.
What kinds of open-source software make personal deployment easy?
Two families dominate. Desktop runners bundle model download, quantized loading, and a chat interface for quick personal trials. Inference server frameworks expose the model behind a standard API for your own programs or a shared team service. The first is faster to start, the second is more flexible.
Does 4-bit quantization noticeably hurt answer quality?
It does, but the damage depends on the task. Factual Q&A and rewriting usually show little difference, while math, multi-step planning, and strict structured output expose it first. If those matter to you, pick a larger model at conservative quantization rather than the smallest one that loads.
What to do next
Measure your VRAM, pick the largest model that runs without offloading, and time one fixed prompt. Those three steps give you your own reference point instead of someone else's build list. If you mainly want a ready-made chat experience, a hosted assistant such as DeepSeek or ChatGPT covers most daily work, and a platform like Dify helps when you need to wire a model into your own system. The two routes are not in conflict; many people run a small local model for sensitive data and send harder tasks to the cloud.
