Qwen3.8-27B API, 140 tok/s on one GPU

infrastructure / Show HN
35
SCORE

A hosted API providing access to the Qwen3.8-27B large language model. The service delivers inference at 140 tokens per second on a single GPU. Users can access the API after logging in through the provided endpoint.

Sources (1)

Score Breakdown

Traction
raw 1.00 · weight 35%
3.5pts
Novelty
0 days old · weight 20%
20.0pts
Source diversity
1 source · weight 10%
3.3pts
AI quality
raw 22.00 · weight 35%
7.7pts

Minimal differentiation—just another hosted LLM API with a speed claim, zero traction, and no clear unique value beyond commodity inference.

Final Score35/100

Similar Products

Marzban-Node
INFRASTRUCTURE · GITHUB
84
SCORE
K3Flight
INFRASTRUCTURE · GITHUB
81
SCORE
Marzban-Panel
INFRASTRUCTURE · GITHUB
80
SCORE
3x-ui-multi
INFRASTRUCTURE · GITHUB
80
SCORE
Qwen-MM-Plugins
AI-TOOLS · GITHUB
80
SCORE
ZEUS-PANEL
INFRASTRUCTURE · GITHUB
79
SCORE
SUPPORT THE PROJECT →
← ALL PRODUCTS© 2026 MARKETHUNT