Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

developer-tools / Show HN
69
SCORE

Run Qwen3.8-Flash-Next, a 104GB model, on 48GB Mac hardware with inference speeds around 12 tokens per second. Implements slot-based streaming for efficient large language model execution on resource-constrained devices.

Sources (1)

Score Breakdown

Traction
raw 115.00 · weight 35%
24.1pts
Novelty
9.2 days old · weight 20%
19.5pts
Source diversity
1 source · weight 10%
3.3pts
AI quality
raw 62.00 · weight 35%
21.7pts

Technically impressive optimization achieving 12 tok/s on constrained hardware, but niche use case with unclear practical value proposition and minimal adoption signals.

Final Score69/100

Similar Products

anti-slop
DEVELOPER-TOOLS · GITHUB
86
SCORE
devspace
DEVELOPER-TOOLS · GITHUB
84
SCORE
eve
DEVELOPER-TOOLS · GITHUB
84
SCORE
kage
DEVELOPER-TOOLS · GITHUB
84
SCORE
openworker
DEVELOPER-TOOLS · GITHUB
84
SCORE
img2threejs
DEVELOPER-TOOLS · GITHUB
84
SCORE
SUPPORT THE PROJECT →
← ALL PRODUCTS© 2026 MARKETHUNT