Llama.cpp fork with 2-4x multiGPU speed for MoE models bigger than VRAM
47
SCORE
A fork of llama.cpp optimized for multi-GPU inference with 2-4x speed improvements on Mixture of Experts models. Enables running MoE models larger than individual GPU VRAM by distributing computation across multiple GPUs.
Sources (1)
1 PTS
Score Breakdown
Traction
raw 2.00 · weight 35%
5.6pts
Novelty
0 days old · weight 20%
20.0pts
Source diversity
1 source · weight 10%
3.3pts
AI quality
raw 52.00 · weight 35%
18.2pts
Technically sound optimization for niche use case (MoE multi-GPU inference), but zero traction and unproven claims of 2-4x speedup limit confidence in execution quality.
Final Score47/100
Similar Products
anti-slop
86
SCORE
devspace
84
SCORE
eve
84
SCORE
kage
84
SCORE
openworker
84
SCORE
img2threejs
84
SCORE
← ALL PRODUCTS© 2026 MARKETHUNT