Decentralized Model Inference Benchmarks

Empirical tokens/second throughput and time-to-first-token (TTFT) across multi-GPU mining rigs and distributed coordinator-worker pairs.

Model Parameters Quantization Hardware Cluster Decode Speed TTFT
Qwen 2.5 0.5 Billion Q4_K_M RTX 3080 + Remote EPYC Node 13.5 tok/s 47 ms
Qwen 2.5 7.0 Billion Q4_K_M 5x 8GB Mining Rig (Vulkan) 24.8 tok/s 62 ms
Llama 3.1 8.0 Billion Q4_K_M Single RTX 4090 24GB 88.4 tok/s 18 ms
Qwen 2.5 32.0 Billion Q4_K_M 8x RTX 3070 8GB (64GB Pool) 18.2 tok/s 85 ms
Llama 3.1 70.0 Billion Q4_K_M Distributed Multi-Rig Cluster 11.6 tok/s 120 ms