Served
Requests and tokens delivered · last hour and last 24 hours
Host
CPU and system memory
Host RAM
used GBGPU
SM, memory, power, and interconnect
GPU SM
util %GPU VRAM
used %GPU memory bandwidth
util %GPU power
wattsGPU temp
°CGPU clocks
SM MHzDCGM tensor
pipe active %DCGM DRAM
HBM active %DCGM NVLink
TX RX GB/sDCGM PCIe
TX RX GB/sGPU fleet
Engine
SGLang and vLLM serving
Requests
running queuedRequests / min
completedKV cache
fill %Cache hit
hit %Total tokens/min
cached prefill + prefill + decodeWaiting tokens
uncachedTTFT
queue prefill totalChunked prefill
secondsE2E latency
secondsITL
ms / tokenSpec accept length
tokensSpec accept rate
%HiCache host
fill %Cache source
device host tok/minHiCache traffic
backup load-back evict tok/minRequest shape
prompt uncached generationHTTP errors
4xx+5xx / minKV retracts
retracted reqsDeployment
—Recent requests
anonymized · no promptsPer-request shape and latency from engine logs. No prompts, completions, IPs, or request ids. Out is generated tokens until the request leaves the decode batch (from gen throughput × time / running).