Loading
Loading
Loading
Prices are our ballpark as of 2026-09-27, not a live store. Memory is what the planner uses.
One DGX Spark holds about a 200 billion parameter model after it is compressed. Two is the size most people finish. Four reaches the largest class and usually needs a fast network switch.
16 of 24 machines
Memory
8 GB RAM, shared with the OS
Price
Already owned
What it can run
A 7B Q4, short context. 14B will swap itself to death.
Best for
Seeing what local feels like. Not an agent.
Memory
16 GB RAM, shared with the OS
Price
Already owned, or $500–1,200 used
What it can run
7B–14B Q4. 32B does not belong here.
Best for
Private notes. A first local chat.
Memory
16–36 GB unified
Price
The laptop you travel with
What it can run
MLX quants from 8B up toward 32B if you bought the memory.
Best for
Quiet private chat on battery.
Memory
12 GB VRAM
Price
$180–280 used
What it can run
14B Q4 comfortably. 32B only with a painful offload.
Best for
A used first GPU.
Memory
8 GB VRAM
Price
$250–450 used
What it can run
7B–14B Q4. 32B does not fit.
Best for
A first new card, not an agent box.
Memory
12 GB VRAM
Price
$450–700 used
What it can run
14B easily. A squeezed 32B only if you accept IQ3 and a short context.
Best for
A quiet mid card.
Memory
16 GB VRAM
Price
$800–1,200 used
What it can run
14B with context. 32B Q4 is tight. 70B does not fit.
Best for
A single-GPU coding chat that is not quite a 4090.
Memory
24 GB VRAM
Price
$1,600–2,200
What it can run
Qwen3 32B Q4 with a short context. 70B is a squeeze.
Best for
Private agents: invoices, mail, Slack.
Memory
32 GB VRAM
Price
$2,400–3,400
What it can run
70B-class Q4 with more context than a 4090. Still one model, one box.
Best for
A single-GPU coding model.
Memory
96 GB VRAM
Price
Workstation money, roughly $8–12k
What it can run
70B with a long context, or a 100B-class quant, still on one card.
Best for
A single workstation that should not become a cluster.
Memory
48 GB VRAM combined
Price
$800–1,400 for the pair
What it can run
70B Q4 if the software splits cleanly.
Best for
Tinkerers who like used hardware.
Memory
48 GB VRAM combined
Price
$3,200–4,400
What it can run
70B Q4 with more speed than a 3090 pair.
Best for
A fast private 70B without buying Spark.
Memory
16–64 GB unified
Price
$600–2,500
What it can run
MLX from 8B to a careful 32B, depending on the memory you ordered.
Best for
A quiet always-on desk box.
Memory
64–192 GB unified
Price
$2,000–8,000
What it can run
MLX quants from 32B toward 70B, quietly, if you bought the memory.
Best for
A private desk that should not sound like a rack.
Memory
Up to 128 GB unified
Price
$1,500–2,500 for a mini PC or Framework Desktop
What it can run
Large MoE and 70B-class GGUF. Bandwidth is the limit, not a missing GPU slot.
Best for
One box, lots of memory, no NVIDIA driver story.
Memory
128 GB coherent
Price
$4,700–6,500
What it can run
About a 200B-class quantized model. A Qwen ~35B is the calmer first success.
Best for
One serious local box. Not a cluster yet.