Talking to
–
–
Last clientnone yet
Status
–
Loaded modelnone
Context–
In memory–
CPU
–%
Load 1 / 5 / 15 min–
Cores–
Memory
–GB
Total–
Available–
Live
–tok/s
Requests in flight0
Tokens streaming0
Elapsed–
Totals
–avg tok/s
Requests0
Generated tokens0
Prompt tokens0
Last 3 minutes — CPU % (blue) and tokens in flight (green)
Recent requests
| Time | Client | Endpoint | Model | Prompt tok | Gen tok | Prompt tok/s | Gen tok/s | Total s |
|---|
No requests recorded yet.
Models
Active–
Status–
| Model | Params | Quant | Size | State |
|---|
Download model
Downloads run on the server in the background; the model is imported with a 64K context and appears in the switcher when done.