Loading
Loading
The same short definitions as the info dots. Tap or hover the i, or read the sentence here.
Tokens
A token is a small piece of text the model reads or writes, often a few letters. Prices are usually counted per million tokens.
Context window
The context window is how much text the model can hold in one conversation, including what you typed and what it writes back.
VRAM
VRAM is the memory on a graphics card. A local model has to fit in that memory.
Local vs cloud
Local means the prompt stays on your machine. Cloud means the prompt is sent to the company that hosts the model.
Quantization
A quantized model is a smaller copy that uses less memory. It is usually a bit less precise than the full model.
Harness
The harness is the app that runs the model, such as Ollama or LM Studio. It is not the model itself.
GGUF
GGUF is a file format for a model that runs on your own computer. A Q4 file is a smaller copy of that file.
API key
An API key is a secret the provider checks before it runs a hosted model. Keep it in a private setting, not in a chat or a screenshot.
Skill index
How well the model handles general questions. Higher is stronger. The source field is intelligence.
Coding score
How well the model writes and edits code. Higher is stronger.
Tools score
How well the model uses tools to finish a multi-step job. The source calls this agentic.