GLM.vLLM
GLM foundation-model testing on vLLM + EKS (invite-only)
Sign in
Models
Each model is one Deployment (scale-to-zero). Start = Karpenter provisions a GPU node (~5-9 min), then vLLM loads the weights. Stop = the node is consolidated (~1-2 min), GPU billing ends.
Chat
Warning: each running model uses its own GPU node - comparing N models costs N times as much
Vision
Upload an image (base64) and ask a vision model (GLM-4.5V / 4.6V / 4.6V-Flash).