Open Infra
connecting
cached 5m

Served

Requests and tokens delivered · last hour and last 24 hours

Host

CPU and system memory

CPU

util % load1

Host RAM

used GB

GPU

SM, memory, power, and interconnect

GPU SM

util %

GPU VRAM

used %

GPU memory bandwidth

util %

GPU power

watts

GPU temp

°C

GPU clocks

SM MHz

DCGM tensor

pipe active %

DCGM DRAM

HBM active %

DCGM NVLink

TX RX GB/s

DCGM PCIe

TX RX GB/s

GPU fleet

Engine

SGLang and vLLM serving

Requests

running queued

Requests / min

completed

KV cache

fill %

Cache hit

hit %

Total tokens/min

cached prefill + prefill + decode

Waiting tokens

uncached

TTFT

queue prefill total

Chunked prefill

seconds

E2E latency

seconds

ITL

ms / token

Spec accept length

tokens

Spec accept rate

%

HiCache host

fill %

Cache source

device host tok/min

HiCache traffic

backup load-back evict tok/min

Request shape

prompt uncached generation

HTTP errors

4xx+5xx / min

KV retracts

retracted reqs

Deployment

Recent requests

anonymized · no prompts

Per-request shape and latency from engine logs. No prompts, completions, IPs, or request ids. Out is generated tokens until the request leaves the decode batch (from gen throughput × time / running).