Ternary-Bonsai-2-27B is the clearest verified local GGUF text-model pick: its smallest published language file is 5.95 GB, needs about 6.8 GB before a modest context cache, and carries an Apache-2.0 licence. It requires Prism ML’s custom llama.cpp build.
At a glance
- Ternary-Bonsai-2-27B packages a 27.36B-parameter model into a 5.95 GB PTQ1_0 GGUF file.
- Our calculation puts the practical starting memory budget for that smallest Ternary-Bonsai file at 6.8 GB, before longer prompts expand the cache.
- XDA’s 20 July 2026 local-model report identifies Qwen 3.5’s 0.8B, 2B, 4B and 9B variants as small-device options, but does not verify a GGUF download for each.
- MiniMax M3’s community licence requires commercial users to display attribution and imposes notice or approval terms tied to annual revenue.
There is an important distinction here. “Open source” is often used for downloadable model weights, even when the licence limits commercial use, branding, redistribution or deployment. BentoML’s 16 June 2026 guide calls this difference “open weights” versus fully open source. For a local user, the practical question is simpler: can the file run on your machine, and may you use it for your purpose
Which local GGUF model fits your hardware?
For a laptop, desktop, or single consumer GPU with roughly 8 GB available for model weights, Ternary-Bonsai-2-27B’s PTQ1_0 file is the verified option in this shortlist. Prism ML lists the model as text generation with GGUF support, a 262K-token context ceiling, and Apache-2.0 licensing.
The unusual part is its ternary storage. Ternary weights use three values rather than conventional higher-precision values, reducing the space occupied by the model. The file remains a 27B-class language model, but the published PTQ1_0 pack is far smaller than its roughly 54 GB FP16 reference file.
| Hardware goal | Published GGUF file | File size | Starting memory budget | What it means |
|---|---|---|---|---|
| Smallest practical local setup | PTQ1_0 | 5.9 GB | 6.8 GB | Best fit when memory is the main limit. |
| More storage, different speed trade-off | PQ2_0 | 7.2 GB | 8.3 GB | Needs more headroom but uses a different packing method. |
| Reference-quality weights | F16 | 53.8 GB | 61.9 GB | Workstation territory rather than an ordinary laptop. |
This is a planning estimate, not a hardware benchmark. The published files and pack descriptions are on the Ternary-Bonsai model page.

Do not treat the budget as a promise of a smooth experience. Context is the text the model keeps available while it answers, and longer chats consume more memory. The model’s own page says PTQ1_0 is the tighter-memory choice, while PQ2_0 can be faster on some newer accelerators. That makes the smaller file the sensible first download for an 8 GB target.
What shipped versus what was promised?
The local-LLM market is full of claims about huge context windows and model efficiency. This dated timeline separates broad architectural promises from files and instructions a reader can actually use.
-
11 June 2026
MiniMax’s Sparse Attention paper described a 109B-parameter test model that cut per-token attention compute by 28.4 times at a one-million-token context. The paper reported 14.2-times faster prompt processing and 7.6-times faster decoding on H800 hardware. Those are research results, not a consumer-laptop promise. -
12 June 2026
The revised MiniMax paper retained its central claim: sparse attention can make very long contexts less expensive than conventional attention. It did not establish a public GGUF package or a verified local-memory figure for MiniMax M3. -
16 June 2026
BentoML’s roundup described MiniMax M3 as a 428B total-parameter mixture-of-experts model with 23B active parameters per token. That framing explains its ambition, but it is not a practical recommendation for ordinary local hardware. -
20 July 2026
XDA’s hands-on shortlist shifted the focus to small models. It highlighted Gemma 4 E2B and Qwen 3.5 models from 0.8B to 9B parameters for phones, Chromebooks and modest PCs. The report is useful hardware guidance, but it is not a licence record or a GGUF file catalogue. -
21 September 2026
Awesome Local LLMs listed more than 10,000 related repositories and tracked local runtimes alongside models. That scale is a reason to start with a specific hardware budget and licence check, rather than choosing the most-starred project.
The shipped item that clears the practical test in this guide is Ternary-Bonsai’s downloadable GGUF package, including two low-bit language-model files and optional vision files. Its limitation is equally concrete: Prism ML’s own runtime fork is part of the deployment requirement.
How to run open-source LLM models locally?
Run a local model by choosing a compatible file, installing a runtime, then keeping the model and its prompts on your machine. llama.cpp’s documentation says it requires GGUF model files and can download compatible files from Hugging Face through its command-line interface.
- Check available RAM or GPU VRAM, then reserve room for your operating system and context cache.
- Pick a GGUF file that fits the budget, not merely the parameter count on a leaderboard.
- Install the runtime named by the model publisher. For Ternary-Bonsai, that means Prism ML’s compatible llama.cpp fork.
- Start with a short context and a small test prompt. Increase context only after the model loads reliably.
- Read the model licence before using the output in a paid product or hosted service.
For the standard GGUF path, a local command can look like this:
llama cli -hf publisher/model:quantizationThat command pattern is useful only when the runtime supports the model’s tensor types. Ternary-Bonsai is the exception that proves the rule: its published instructions require the project-specific fork. Installing a familiar local app first and assuming compatibility later is a common route to cryptic load errors.
What is the smallest LLM model I can run locally?
The smallest text model named in the current small-model reporting here is Qwen 3.5 0.8B, according to XDA’s July report. A “B” means billions of parameters, or learned values that shape a model’s responses. Fewer parameters usually means a smaller download and lower memory demand.
Smallest is not automatically best. A sub-1B model can be appropriate for brief offline help, simple extraction, classification and lightweight structured tasks. It is less likely to be satisfying for long coding sessions, nuanced writing, multi-step reasoning or large document analysis. The right buying rule is to choose the smallest model that completes your real task, then move up only when it fails.
There is also a category trap. Nvidia’s Nemotron-3-Diarization is extremely small at 0.1B parameters, but diarization means identifying who spoke when in audio. It is not a general chat LLM, so it should not be compared with a local text assistant.
Are there any open-source LLM models?
Yes, but downloadable does not always mean unrestricted. Instaclustr’s 2026 guide describes the core distinction: some projects publish weights, code and permissive licences, while others publish usable weights with additional legal conditions. Both can run locally, but they offer different rights.
Ternary-Bonsai lists Apache-2.0, a familiar permissive software licence. MiniMax M3 uses a community licence with different terms. MiniMax’s published licence requires commercial users to display “Built with MiniMax M3”; it also requires a one-time notice below $20 million in annual revenue and prior written authorization above that threshold.
That difference matters more than a benchmark rank for a business. A hobbyist can reasonably focus on fit and output quality. A company should keep a copy of the exact licence version, identify whether its use is commercial, and ask counsel to review terms that require notice, attribution or approval.
What is the best open-source LLM I can run locally?
The best local model is the one that fits your hardware, runtime and licence needs without turning routine work into a waiting game. For a verified GGUF text-model download near an 8 GB memory target, Ternary-Bonsai-2-27B is the strongest specific recommendation in this guide because its file sizes, licence and runtime condition are published together.
For older or lower-power machines, use the small-model direction rather than forcing a 27B-class model into swap memory. A June consumer-hardware test found that 7B to 12B models were the workable middle ground on a PC with 12 GB of VRAM, while larger quantized models became much slower. That is an experience report, not a universal benchmark, but it reflects the decision users actually face.
For coding, ask whether the model has code-specific training and whether it follows structured output. For private document work, prioritize local execution and a context length you can afford. For travel or offline use, a 0.8B to 4B model can be more valuable than a larger model that stays at home.
What we could not verify?
Public material does not establish a single, comparable local benchmark for every model mentioned in current rankings. It also does not establish that every small model discussed in hands-on coverage has a maintained GGUF download, a current official quantization, or a licence unchanged from its original release.
MiniMax could settle the remaining M3 questions by publishing an official GGUF release, exact local-memory guidance and a support matrix for consumer runtimes. Model publishers could make local deployment clearer by listing the required runner version, the tested context length and the licence terms beside every downloadable file.
Sources
- Prism ML: Ternary-Bonsai-2-27B GGUF model card
- llama.cpp: obtaining and running GGUF models
- MiniMax Sparse Attention paper
- BentoML: The Best Open-Source LLMs in 2026
- XDA: tiny local LLM testing report
- Awesome Local LLMs repository
- MiniMax Community License for M3
- Instaclustr: Top open-source LLMs for 2026
- Mayhem Code: consumer-hardware local LLM testing

