REFFT AI Store

A native serving and inference tool.

Add model weights, and you have the AI assistant you need.

8 of 8 models
GLM-4.6V-FlashMeasured

Instruction-tuned multimodal GLM-4.6V-Flash language model for on-device text generation.

4.50tok/s
TPS / Decode
29.8tok/s
Prefill
128
Context
Text-to-TextQualcomm Hexagon NPUNVIDIA GPUAMD GPUApple NPUw4QCS9075
LFM2.5-2.6BMeasured

LiquidAI LFM2.5-2.6B hybrid dense decoder for on-device text generation.

31.8tok/s
TPS / Decode
150tok/s
Prefill
1,024
Context
Text-to-TextQualcomm Hexagon NPUw4SM8850+1
LFM2.5-8B-A1BMeasured

LiquidAI LFM2.5-8B-A1B hybrid MoE language model for on-device text generation.

25.3tok/s
TPS / Decode
23.3tok/s
Prefill
512
Context
Text-to-TextQualcomm Hexagon NPUw4QCS9075
LFM2.5-VL-3BMeasured

LiquidAI LFM2.5-VL-3B hybrid decoder + SigLIP2 (text-first Hexagon pack).

33.4tok/s
TPS / Decode
159tok/s
Prefill
1,024
Context
Image-Text-to-TextQualcomm Hexagon NPUw4SM8850+1
Qwen3-0.6BMeasured

Qwen3-0.6B dense decoder language model for on-device text generation.

76.4tok/s
TPS / Decode
860tok/s
Prefill
1,024
Context
Text-to-TextQualcomm Hexagon NPUw4QCS9075
Qwen3-1.7BMeasured

Qwen3-1.7B dense decoder language model for on-device text generation.

38.5tok/s
TPS / Decode
668tok/s
Prefill
1,024
Context
Text-to-TextQualcomm Hexagon NPUw4QCS9075
Qwen3-4B-Instruct-2507Measured

Qwen3-4B-Instruct-2507 dense decoder language model for on-device text generation.

22.8tok/s
TPS / Decode
545tok/s
Prefill
1,024
Context
Text-to-TextQualcomm Hexagon NPUw4QCS9075
SmolVLM2-500M-Video-InstructMeasured

Compact vision-language model that captions images and answers questions about them on-device, with text, image and video-frame input.

99.8tok/s
TPS / Decode
1,725tok/s
Prefill
1,024
Context
Image-Text-to-TextQualcomm Hexagon NPUNVIDIA GPUAMD GPUApple NPUw4SM8850+1

Catalog refreshed from GitHub every 300s