Explainers
Local LLM runners explained: running models on your computer
What running an AI model locally means, how hardware limits it, when it makes sense, and how local runners such as Ollama compare with hosted chatbots.
A local LLM runner loads an open AI model onto your own computer and generates text there, instead of sending each request to a company's servers. The model is a set of files, and the runner is the software that reads those files and does the computation. Hardware, especially memory, decides which models are practical. Local runners trade the convenience and the frontier-level models of hosted chatbots for control over where the model runs.
The best local LLM runners ranking compares six tools. Its scores are editorial judgments on setup, interface, license and control, not measurements of speed or answer quality, since both depend on your hardware and the model you choose.
What a runner does
A runner has three jobs. It loads a model file into memory, it runs the model to produce text, and it gives you a way to talk to it, usually a chat window, a command line or a local server. Some tools do all three. Others are only the engine.
The six tools in the ranking fall into roughly three groups:
- Runtimes. Ollama wraps models in a simple install and a model library, and it runs on macOS, Linux and Windows. llama.cpp is the engine many other tools build on. It is a C and C++ project with a server and a library API, and it needs command-line comfort.
- Desktop chat apps. Jan and GPT4All give you a chat window without a terminal. GPT4All also offers chat over local documents.
- Packaging and servers. llamafile bundles a model and its runtime into one executable file. LocalAI runs as a self-hosted server for text, image, voice and video models.
Hardware limits, in general terms
The main limit is memory. A model has to fit in the memory the runner can use, either system RAM or the memory on a graphics processor. Larger models generally need more memory. Many runners use compressed versions of models, usually called quantized, which need less memory at some cost in output quality. The exact size you can run depends on the model and the file format, so check the requirements of the specific model you want.
Speed depends on the processor and the memory. A graphics processor that the runner supports can speed up generation a lot, and a computer without one can still run smaller models more slowly. Ollama, Jan and GPT4All all state that their speed depends on your hardware, and that is the honest answer for every runner in this group.
Some practical details also matter. llamafile's listing notes that Windows cannot run executables larger than 4 GB, which rules out larger models on that system. A single model file can be very large on disk as well as in memory, so check free storage too.
When running locally makes sense
Local runners fit a few kinds of work well:
- Offline use. Jan runs models offline on your own computer, so the work does not need a connection once the model is installed.
- Sensitive text you prefer not to send to a vendor. With a fully local setup, the prompts stay on your device. Local does not mean secure on its own, so you still need to protect the computer and any files the model can read.
- Development and testing. llama.cpp and LocalAI give developers a server or a library they can build on, which is useful for trying an app against a model without paying per request.
- Predictable costs. Local models have no per-message bill, but you pay for the hardware, the electricity and your time.
Local runners fit less well when you need the most capable models, live web search or current information, or a setup that many people share without an administrator. A hosted chatbot is usually simpler for those.
Local runners against hosted chatbots
The two approaches answer different needs. The table compares them on the questions that usually decide it.
| Question | Local runner | Hosted chatbot |
|---|---|---|
| Where the model runs | Your computer or your own server | The vendor's servers |
| Which models | Open models you download, with licenses that vary | The models the vendor offers |
| Web search and files | Depends on the app | Often built in, such as web search on the free plans of ChatGPT and Claude |
| Cost | Hardware, electricity and setup time | Subscription or usage fees |
| Speed | Set by your hardware | Set by the vendor and your plan |
| Updates | You download new models | The vendor updates the service |
| Data | Stays on your device unless you connect a cloud model | Sent to the vendor under its terms |
If you want to compare the hosted side in detail, the best AI chatbots ranking covers the main assistants and their plans. Ollama and Jan can also connect to cloud models, so the line between the two approaches is not always sharp. Check where each request goes before you assume it stays local.
How to pick a runner
Pick by how you will use the model rather than by its name:
- Comfortable with a terminal and want a simple install: Ollama.
- Want the engine itself for development: llama.cpp.
- Want a desktop chat window: Jan or GPT4All.
- Want to share one model as a single file: llamafile.
- Want a self-hosted server for several model types: LocalAI.
Check the license of each model as well as the runner. The runners in this group are mostly MIT or Apache 2.0, which allow commercial use, but a model can carry its own terms.
What to do next
Install one runner, download a small model that your computer can handle comfortably, and try the tasks you care about. Compare the answers with a hosted assistant on the same prompts, and note how long each takes on your machine. That test will show you more about fit than any general claim, including the ones in this article.
Mentioned in this article
OllamaRun open AI models locally with a simple command-line install
llama.cppC and C++ engine for running local LLMs, with a server and tools
JanOpen-source desktop app for running AI models on your computer
GPT4AllDesktop app for running local open-source models on Windows, macOS and Linux
llamafileSingle-file executables that bundle a language model and its runtime
LocalAISelf-hosted open-source AI server for text, image, voice and video models