Llama.cpp fork with 2-4x multiGPU speed for MoE models bigger than VRAM

developer-tools / Show HN
47
SCORE

A fork of llama.cpp optimized for multi-GPU inference with 2-4x speed improvements on Mixture of Experts models. Enables running MoE models larger than individual GPU VRAM by distributing computation across multiple GPUs.

Sources (1)

Score Breakdown

Traction
raw 2.00 · weight 35%
5.6pts
Novelty
0 days old · weight 20%
20.0pts
Source diversity
1 source · weight 10%
3.3pts
AI quality
raw 52.00 · weight 35%
18.2pts

Technically sound optimization for niche use case (MoE multi-GPU inference), but zero traction and unproven claims of 2-4x speedup limit confidence in execution quality.

Final Score47/100

Similar Products

anti-slop
DEVELOPER-TOOLS · GITHUB
86
SCORE
devspace
DEVELOPER-TOOLS · GITHUB
84
SCORE
eve
DEVELOPER-TOOLS · GITHUB
84
SCORE
kage
DEVELOPER-TOOLS · GITHUB
84
SCORE
openworker
DEVELOPER-TOOLS · GITHUB
84
SCORE
img2threejs
DEVELOPER-TOOLS · GITHUB
84
SCORE
SUPPORT THE PROJECT →
← ALL PRODUCTS© 2026 MARKETHUNT