Skip to content

Commit 9262932

Browse files
authored
Merge pull request #131 from zhongkaifu/feature/merge_two_tables_into_single_table
merge two tables into a single table
2 parents 30e711f + db68acb commit 9262932

1 file changed

Lines changed: 30 additions & 58 deletions

File tree

docs/moe_cpu_offload_benchmark.md

Lines changed: 30 additions & 58 deletions
Original file line numberDiff line numberDiff line change
@@ -109,62 +109,40 @@ pass 20/20 on `ggml_cuda` (±offload), `ggml_vulkan` and `cuda`.
109109

110110
## Results by model
111111

112-
Ratios are TensorSharp / llama.cpp: **>1.0x means TensorSharp is faster**, and for
113-
VRAM **>1.0x means TensorSharp is heavier**.
112+
Each row is one offload depth, with TensorSharp, llama.cpp and the ratio between
113+
them side by side for every metric. Ratios are TensorSharp / llama.cpp: **>1.0x
114+
means TensorSharp is faster**, and for VRAM **>1.0x means TensorSharp is
115+
heavier**.
114116

115117
### Gemma 4 26B-A4B it (UD-IQ4_XS, 30 MoE layers)
116118

117-
| `--n-cpu-moe` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
118-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
119-
| 0 _(baseline)_ | 16,822 | 11,173 | 11,274 | 161.4 | 14,602 | 10,843 | 10,628 | 206.7 |
120-
| 8 | 15,724 | 7,063 | 6,500 | 80.2 | 11,874 | 1,459 | 1,459 | 32.7 |
121-
| 16 | 14,128 | 4,183 | 4,888 | 54.5 | 9,122 | 833 | 854 | 21.9 |
122-
| 24 | 12,346 | 3,500 | 3,958 | 49.1 | 6,368 | 667 | 689 | 16.7 |
123-
| 30 _(`--cpu-moe`)_ | 11,038 | 3,035 | 3,072 | 39.7 | 4,134 | 543 | 495 | 12.9 |
124-
125-
| `--n-cpu-moe` | VRAM | pp4096 | pp8192 | tg128 |
126-
|---|---:|---:|---:|---:|
127-
| 0 | 1.15x | **1.03x** | **1.06x** | 0.78x |
128-
| 8 | 1.32x | **4.84x** | **4.46x** | **2.45x** |
129-
| 16 | 1.55x | **5.02x** | **5.72x** | **2.49x** |
130-
| 24 | 1.94x | **5.25x** | **5.74x** | **2.93x** |
131-
| 30 | 2.67x | **5.59x** | **6.21x** | **3.07x** |
119+
| `--n-cpu-moe` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
120+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
121+
| 0 _(baseline)_ | 16,822 | 14,602 | 1.15x | 11,173 | 10,843 | **1.03x** | 11,274 | 10,628 | **1.06x** | 161.4 | 206.7 | 0.78x |
122+
| 8 | 15,724 | 11,874 | 1.32x | 7,063 | 1,459 | **4.84x** | 6,500 | 1,459 | **4.46x** | 80.2 | 32.7 | **2.45x** |
123+
| 16 | 14,128 | 9,122 | 1.55x | 4,183 | 833 | **5.02x** | 4,888 | 854 | **5.72x** | 54.5 | 21.9 | **2.49x** |
124+
| 24 | 12,346 | 6,368 | 1.94x | 3,500 | 667 | **5.25x** | 3,958 | 689 | **5.74x** | 49.1 | 16.7 | **2.93x** |
125+
| 30 _(`--cpu-moe`)_ | 11,038 | 4,134 | 2.67x | 3,035 | 543 | **5.59x** | 3,072 | 495 | **6.21x** | 39.7 | 12.9 | **3.07x** |
132126

133127
### Qwen 3.5 35B-A3B (UD-IQ4_XS, 48 MoE layers)
134128

135-
| `--n-cpu-moe` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
136-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
137-
| 0 _(baseline)_ | 19,862 | 9,538 | 9,405 | 160.0 | 17,522 | 8,149 | 8,073 | 228.4 |
138-
| 12 | 18,148 | 6,755 | 6,648 | 75.4 | 13,282 | 988 | 954 | 27.5 |
139-
| 24 | 15,414 | 4,412 | 5,259 | 52.3 | 9,010 | 498 | 484 | 15.8 |
140-
| 36 | 12,684 | 3,772 | 4,223 | 50.7 | 4,738 | 523 | 517 | 11.3 |
141-
| 48 _(`--cpu-moe`)_ | 11,606 | 3,917 | 3,709 | 38.6 | 3,314 | 477 | 457 | 10.2 |
142-
143-
| `--n-cpu-moe` | VRAM | pp4096 | pp8192 | tg128 |
144-
|---|---:|---:|---:|---:|
145-
| 0 | 1.13x | **1.17x** | **1.16x** | 0.70x |
146-
| 12 | 1.37x | **6.84x** | **6.97x** | **2.74x** |
147-
| 24 | 1.71x | **8.85x** | **10.86x** | **3.31x** |
148-
| 36 | 2.68x | **7.21x** | **8.17x** | **4.50x** |
149-
| 48 | 3.50x | **8.21x** | **8.11x** | **3.77x** |
129+
| `--n-cpu-moe` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
130+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
131+
| 0 _(baseline)_ | 19,862 | 17,522 | 1.13x | 9,538 | 8,149 | **1.17x** | 9,405 | 8,073 | **1.16x** | 160.0 | 228.4 | 0.70x |
132+
| 12 | 18,148 | 13,282 | 1.37x | 6,755 | 988 | **6.84x** | 6,648 | 954 | **6.97x** | 75.4 | 27.5 | **2.74x** |
133+
| 24 | 15,414 | 9,010 | 1.71x | 4,412 | 498 | **8.85x** | 5,259 | 484 | **10.86x** | 52.3 | 15.8 | **3.31x** |
134+
| 36 | 12,684 | 4,738 | 2.68x | 3,772 | 523 | **7.21x** | 4,223 | 517 | **8.17x** | 50.7 | 11.3 | **4.50x** |
135+
| 48 _(`--cpu-moe`)_ | 11,606 | 3,314 | 3.50x | 3,917 | 477 | **8.21x** | 3,709 | 457 | **8.11x** | 38.6 | 10.2 | **3.77x** |
150136

151137
### GPT-OSS 20B (Q8_0 / MXFP4, 24 MoE layers)
152138

153-
| `--n-cpu-moe` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
154-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
155-
| 0 _(baseline)_ | 13,186 | 13,964 | 12,925 | 212.8 | 12,204 | 17,856 | 17,642 | 344.2 |
156-
| 6 | 11,560 | 8,975 | 7,617 | 85.8 | 9,812 | 1,747 | 1,666 | 32.2 |
157-
| 12 | 9,378 | 6,470 | 6,394 | 51.7 | 7,386 | 1,176 | 1,188 | 18.3 |
158-
| 18 | 7,192 | 4,315 | 4,393 | 30.7 | 4,962 | 807 | 751 | 12.1 |
159-
| 24 _(`--cpu-moe`)_ | 4,762 | 4,277 | 3,798 | 27.7 | 2,536 | 568 | 548 | 9.4 |
160-
161-
| `--n-cpu-moe` | VRAM | pp4096 | pp8192 | tg128 |
162-
|---|---:|---:|---:|---:|
163-
| 0 | 1.08x | 0.78x | 0.73x | 0.62x |
164-
| 6 | 1.18x | **5.14x** | **4.57x** | **2.67x** |
165-
| 12 | 1.27x | **5.50x** | **5.38x** | **2.83x** |
166-
| 18 | 1.45x | **5.35x** | **5.85x** | **2.54x** |
167-
| 24 | 1.88x | **7.53x** | **6.93x** | **2.95x** |
139+
| `--n-cpu-moe` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
140+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
141+
| 0 _(baseline)_ | 13,186 | 12,204 | 1.08x | 13,964 | 17,856 | 0.78x | 12,925 | 17,642 | 0.73x | 212.8 | 344.2 | 0.62x |
142+
| 6 | 11,560 | 9,812 | 1.18x | 8,975 | 1,747 | **5.14x** | 7,617 | 1,666 | **4.57x** | 85.8 | 32.2 | **2.67x** |
143+
| 12 | 9,378 | 7,386 | 1.27x | 6,470 | 1,176 | **5.50x** | 6,394 | 1,188 | **5.38x** | 51.7 | 18.3 | **2.83x** |
144+
| 18 | 7,192 | 4,962 | 1.45x | 4,315 | 807 | **5.35x** | 4,393 | 751 | **5.85x** | 30.7 | 12.1 | **2.54x** |
145+
| 24 _(`--cpu-moe`)_ | 4,762 | 2,536 | 1.88x | 4,277 | 568 | **7.53x** | 3,798 | 548 | **6.93x** | 27.7 | 9.4 | **2.95x** |
168146

169147
> GPT-OSS is the one model still behind on the resident path at these prompt
170148
> lengths (0.73–0.78x). Moving to 4,096/8,192-token prompts is what exposed it:
@@ -183,17 +161,11 @@ DeepSeek V4 does not use the host-MoE seam: it hands its offloaded layers to
183161
`ggml_backend_sched`, which streams the weights back to the GPU for a
184162
prefill-sized batch exactly the way llama.cpp does.
185163

186-
| `--n-cpu-moe` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
187-
|---|---:|---:|---:|---:|---:|---:|---:|---:|
188-
| 0 _(baseline, both GPUs)_ | 169,132 | 3,448 | 4,387 | 51.1 | 155,608 | 2,398 | 2,232 | 49.6 |
189-
| 12 | 131,818 | 392 | 428 | 10.3 | 117,150 | 126 | 124 | 13.7 |
190-
| 24 | 79,742 | 218 | 236 | 5.3 | 78,954 | 64 | 63 | 7.2 |
191-
192-
| `--n-cpu-moe` | VRAM | pp4096 | pp8192 | tg128 |
193-
|---|---:|---:|---:|---:|
194-
| 0 | 1.09x | **1.44x** | **1.97x** | **1.03x** |
195-
| 12 | 1.13x | **3.11x** | **3.46x** | 0.75x |
196-
| 24 | 1.01x | **3.42x** | **3.72x** | 0.74x |
164+
| `--n-cpu-moe` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
165+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
166+
| 0 _(baseline, both GPUs)_ | 169,132 | 155,608 | 1.09x | 3,448 | 2,398 | **1.44x** | 4,387 | 2,232 | **1.97x** | 51.1 | 49.6 | **1.03x** |
167+
| 12 | 131,818 | 117,150 | 1.13x | 392 | 126 | **3.11x** | 428 | 124 | **3.46x** | 10.3 | 13.7 | 0.75x |
168+
| 24 | 79,742 | 78,954 | 1.01x | 218 | 64 | **3.42x** | 236 | 63 | **3.72x** | 5.3 | 7.2 | 0.74x |
197169

198170
DeepSeek V4 remains the one model where llama.cpp's offloaded *decode* is ahead
199171
(0.74–0.75x): both engines run that matmul through the same ggml CPU backend, and

0 commit comments

Comments
 (0)