REFFT AI Store

Qwen3-4B-Instruct-2507

Qwen3-4B-Instruct-2507 dense decoder language model for on-device text generation.

Qwen3-4B-Instruct-2507 is Alibaba's 36-layer dense decoder (GQA attention, SwiGLU FFN). The hexagon package is a host VISA client (qwen3_4b) talking to refft-hexagon vm on the board. Modeling is built from refft.hexagon.ops.

Model performance

Qwen3-4B-Instruct-2507 on Hexagon HTP via refft-hexagon cli --model (1×HTP and 2×HTP as two devices).

Dragonwing IQ-9075 (QCS9075) 2×HTP
SoC QCS9075 · Ubuntu 24.04.3 LTS (aarch64) · Hexagon v73 · w4 · runtime 0.6.5.dev682+g9833e0505 · measured 2026-08-13
Measured · QCS9075
22.8tok/s
TPS / Decode speed
113 ms
TTFT
545tok/s
Prefill speed

Install and run

Install
curl -fsSL https://raw.githubusercontent.com/refinefuture-ai/refft.cpp/main/refft-hexagon/install.sh | sh
Run
~/.local/share/refft-hexagon/bin/refft-hexagon cli --model ./Qwen3-4B-Instruct-2507-w4.refft --backend hexagon --prompt "Who are you?" --max_new_tokens 128
Serving, environment variables and full instructions

Tags

llmgenerative-aiquantizedqwenqwen3