Ollama Monitor

connecting…

Talking to

Last clientnone yet

Status

Loaded modelnone
Context
In memory

CPU

%
Load 1 / 5 / 15 min
Cores

Memory

GB
Total
Available

Live

tok/s
Requests in flight0
Tokens streaming0
Elapsed

Totals

avg tok/s
Requests0
Generated tokens0
Prompt tokens0

Last 3 minutes — CPU % (blue) and tokens in flight (green)

Recent requests

TimeClientEndpointModelPrompt tokGen tokPrompt tok/sGen tok/sTotal s
No requests recorded yet.

Models

Active
Status
ModelParamsQuantSizeState

Download model

Downloads run on the server in the background; the model is imported with a 64K context and appears in the switcher when done.

Quick test

streams a short reply through the same endpoint VS Code uses