Loading
Loading
The portable GGUF runtime.
The portable GGUF runtime. It is the recipe for laptops, used NVIDIA cards, Macs, and Ryzen AI Max boxes when you do not want a Python server.
No install command is stored for this one.
The README lists llama.app, Docker, release binaries, and a build guide. No single install command is stored here.
b11371 · Oct 3
Update for how model files are packed
2026-09-22
A desktop coder, not the V4 Spark card.
Distilled reasoning size. A Spark or a pair, not a 4090.
Four Sparks. Not the two-node V4 Flash card.
An older small Gemma. Still a floor, not a coder.
A 16 GB laptop can hold it. A 4090 is wasted on it.
The honest small-laptop ceiling for Gemma.
34 more are on the models index.