Loading
Loading
Loading
Prices are our ballpark as of 2026-09-27, not a live store. Memory is what the planner uses.
One DGX Spark holds about a 200 billion parameter model after it is compressed. Two is the size most people finish. Four reaches the largest class and usually needs a fast network switch.
8 of 24 machines
Memory
8 GB RAM, shared with the OS
Price
Already owned
What it can run
A 7B Q4, short context. 14B will swap itself to death.
Best for
Seeing what local feels like. Not an agent.
Memory
16 GB RAM, shared with the OS
Price
Already owned, or $500–1,200 used
What it can run
7B–14B Q4. 32B does not belong here.
Best for
Private notes. A first local chat.
Memory
16–36 GB unified
Price
The laptop you travel with
What it can run
MLX quants from 8B up toward 32B if you bought the memory.
Best for
Quiet private chat on battery.
Memory
12 GB VRAM
Price
$180–280 used
What it can run
14B Q4 comfortably. 32B only with a painful offload.
Best for
A used first GPU.
Memory
8 GB VRAM
Price
$250–450 used
What it can run
7B–14B Q4. 32B does not fit.
Best for
A first new card, not an agent box.
Memory
12 GB VRAM
Price
$450–700 used
What it can run
14B easily. A squeezed 32B only if you accept IQ3 and a short context.
Best for
A quiet mid card.
Memory
16 GB VRAM
Price
$800–1,200 used
What it can run
14B with context. 32B Q4 is tight. 70B does not fit.
Best for
A single-GPU coding chat that is not quite a 4090.
Memory
24 GB VRAM
Price
$1,600–2,200
What it can run
Qwen3 32B Q4 with a short context. 70B is a squeeze.
Best for
Private agents: invoices, mail, Slack.