REFFT AI Store

GLM-4.6V-Flash

Instruction-tuned multimodal GLM-4.6V-Flash language model for on-device text generation.

GLM-4.6V-Flash is ZhipuAI's compact multimodal language model. This AI Hub entry focuses on the text decoder as a refft serving pack (W4DA16-PERF, multi-part paged weights). Vision paths and compute backends depend on the hardware package (see hexagon/, cuda/, mlx/, amd/).

Model performance

GLM-4.6V-Flash text decoder on Hexagon HTP via refft-hexagon (W4DA16 serving pack).

Dragonwing IQ-9075 (QCS9075) 1×HTP
SoC QCS9075 · Qualcomm Linux Robotics Reference Distro with ROS(SOTA and SELinux enabled) (SELinux-enabled) (OTA-enabled) 2.0 (aarch64) · Hexagon v73 · w4 · runtime v2026.07.26.00 · measured 2026-08-06
Measured · QCS9075
4.50tok/s
TPS / Decode speed
635 ms – 638 ms
TTFT
29.8tok/s
Prefill speed

Install and run

Install
curl -fsSL https://raw.githubusercontent.com/refinefuture-ai/refft.cpp/main/refft-hexagon/install.sh | sh
Run
~/.local/share/refft-hexagon/bin/refft-hexagon cli --model ./GLM-4.6V-Flash-W4.refft --backend hexagon --prompt "Who are you?" --max_new_tokens 128
Serving, environment variables and full instructions

Tags

llmgenerative-aiquantizedglmglm46v
GLM-4.6V-Flash -- REFFT AI Store