Dragonwing IQ-9075 (QCS9075) 2×HTP
22.8tok/s
TPS / Decode speed
113 ms
TTFT
545tok/s
Prefill speed
Qwen3-4B-Instruct-2507 dense decoder language model for on-device text generation.
Qwen3-4B-Instruct-2507 is Alibaba's 36-layer dense decoder (GQA attention, SwiGLU FFN). The hexagon package is a host VISA client (qwen3_4b) talking to refft-hexagon vm on the board. Modeling is built from refft.hexagon.ops.
Qwen3-4B-Instruct-2507 on Hexagon HTP via refft-hexagon cli --model (1×HTP and 2×HTP as two devices).
curl -fsSL https://raw.githubusercontent.com/refinefuture-ai/refft.cpp/main/refft-hexagon/install.sh | sh~/.local/share/refft-hexagon/bin/refft-hexagon cli --model ./Qwen3-4B-Instruct-2507-w4.refft --backend hexagon --prompt "Who are you?" --max_new_tokens 128