Loading
Loading
An OpenAI-shaped API on your own metal.
The serving stack for a card with room: a 5090, a 96 GB workstation GPU, or a pair. Use it when you want an OpenAI-shaped API on your own metal.
uv pip install vllm --torch-backend=auto
From the stable GPU install docs. Other hardware stays on the install page, with no command stored here.
v0.30.0 · Sep 22
New models: DeepSeek-V4.1-Flash with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100, DeepGEMM Mega-mHC, and async Engram prefetch with Engram DP…
2026-09-22
2026-09-09
2026-08-26
2026-08-11
2026-08-10
Four nodes. Two Sparks do not hold these weights.
A heavier quant: trades a little fidelity for a longer context.
Four Sparks and usually a 200GbE switch. A weekend, not an evening.
Measured near 36 tok/s single-stream on two Sparks.
A pair of 3090s or one big card. On a pair the link is the bottleneck.
Served rather than chatted with, when something else calls the model.
5 more are on the models index.