Performance Metrics
Decentralized Model Inference Benchmarks
Empirical tokens/second throughput and time-to-first-token (TTFT) across multi-GPU mining rigs and distributed coordinator-worker pairs.
| Model | Parameters | Quantization | Hardware Cluster | Decode Speed | TTFT |
|---|---|---|---|---|---|
| Qwen 2.5 | 0.5 Billion | Q4_K_M | RTX 3080 + Remote EPYC Node | 13.5 tok/s | 47 ms |
| Qwen 2.5 | 7.0 Billion | Q4_K_M | 5x 8GB Mining Rig (Vulkan) | 24.8 tok/s | 62 ms |
| Llama 3.1 | 8.0 Billion | Q4_K_M | Single RTX 4090 24GB | 88.4 tok/s | 18 ms |
| Qwen 2.5 | 32.0 Billion | Q4_K_M | 8x RTX 3070 8GB (64GB Pool) | 18.2 tok/s | 85 ms |
| Llama 3.1 | 70.0 Billion | Q4_K_M | Distributed Multi-Rig Cluster | 11.6 tok/s | 120 ms |