GLM.vLLM
Running GPU nodes are being billed - press Stop when finished

GLM.vLLM

GLM foundation-model testing on vLLM + EKS (invite-only)

Sign in

Models

Each model is one Deployment (scale-to-zero). Start = Karpenter provisions a GPU node (~5-9 min), then vLLM loads the weights. Stop = the node is consolidated (~1-2 min), GPU billing ends.

Chat

Warning: each running model uses its own GPU node - comparing N models costs N times as much

Vision

Upload an image (base64) and ask a vision model (GLM-4.5V / 4.6V / 4.6V-Flash).

preview