Qwen3.8-27B API, 140 tok/s on one GPU

infrastructure / Show HN
35
SCORE

A hosted API providing access to the Qwen3.8-27B large language model. The service delivers inference at 140 tokens per second on a single GPU. Users can access the API after logging in through the provided endpoint.

Sources (1)

Score Breakdown

Traction
raw 1.00 · weight 35%
3.5pts
Novelty
0 days old · weight 20%
20.0pts
Source diversity
1 source · weight 10%
3.3pts
AI quality
raw 22.00 · weight 35%
7.7pts

Minimal differentiation—just another hosted LLM API with a speed claim, zero traction, and no clear unique value beyond commodity inference.

Final Score35/100

Similar Products

Marzban-Node
INFRASTRUCTURE · GITHUB
84
SCORE
Qwen3.8-27B-GGUF
DEVELOPER-TOOLS · HUGGING FACE
84
SCORE
Qwen3.8-Flash-Next
DEVELOPER-TOOLS · HUGGING FACE
82
SCORE
K3Flight
INFRASTRUCTURE · GITHUB
81
SCORE
Marzban-Panel
INFRASTRUCTURE · GITHUB
80
SCORE
3x-ui-multi
INFRASTRUCTURE · GITHUB
80
SCORE
SUPPORT THE PROJECT →
← ALL PRODUCTS© 2026 MARKETHUNT