Dragonwing IQ-9075 (QCS9075) 1×HTP
4.50tok/s
TPS / Decode speed
635 ms – 638 ms
TTFT
29.8tok/s
Prefill speed
Instruction-tuned multimodal GLM-4.6V-Flash language model for on-device text generation.
GLM-4.6V-Flash is ZhipuAI's compact multimodal language model. This AI Hub entry focuses on the text decoder as a refft serving pack (W4DA16-PERF, multi-part paged weights). Vision paths and compute backends depend on the hardware package (see hexagon/, cuda/, mlx/, amd/).
GLM-4.6V-Flash text decoder on Hexagon HTP via refft-hexagon (W4DA16 serving pack).
curl -fsSL https://raw.githubusercontent.com/refinefuture-ai/refft.cpp/main/refft-hexagon/install.sh | sh~/.local/share/refft-hexagon/bin/refft-hexagon cli --model ./GLM-4.6V-Flash-W4.refft --backend hexagon --prompt "Who are you?" --max_new_tokens 128