@@ -109,62 +109,40 @@ pass 20/20 on `ggml_cuda` (±offload), `ggml_vulkan` and `cuda`.
109109
110110## Results by model
111111
112- Ratios are TensorSharp / llama.cpp: ** >1.0x means TensorSharp is faster** , and for
113- VRAM ** >1.0x means TensorSharp is heavier** .
112+ Each row is one offload depth, with TensorSharp, llama.cpp and the ratio between
113+ them side by side for every metric. Ratios are TensorSharp / llama.cpp: ** >1.0x
114+ means TensorSharp is faster** , and for VRAM ** >1.0x means TensorSharp is
115+ heavier** .
114116
115117### Gemma 4 26B-A4B it (UD-IQ4_XS, 30 MoE layers)
116118
117- | ` --n-cpu-moe ` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
118- | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
119- | 0 _ (baseline)_ | 16,822 | 11,173 | 11,274 | 161.4 | 14,602 | 10,843 | 10,628 | 206.7 |
120- | 8 | 15,724 | 7,063 | 6,500 | 80.2 | 11,874 | 1,459 | 1,459 | 32.7 |
121- | 16 | 14,128 | 4,183 | 4,888 | 54.5 | 9,122 | 833 | 854 | 21.9 |
122- | 24 | 12,346 | 3,500 | 3,958 | 49.1 | 6,368 | 667 | 689 | 16.7 |
123- | 30 _ (` --cpu-moe ` )_ | 11,038 | 3,035 | 3,072 | 39.7 | 4,134 | 543 | 495 | 12.9 |
124-
125- | ` --n-cpu-moe ` | VRAM | pp4096 | pp8192 | tg128 |
126- | ---| ---:| ---:| ---:| ---:|
127- | 0 | 1.15x | ** 1.03x** | ** 1.06x** | 0.78x |
128- | 8 | 1.32x | ** 4.84x** | ** 4.46x** | ** 2.45x** |
129- | 16 | 1.55x | ** 5.02x** | ** 5.72x** | ** 2.49x** |
130- | 24 | 1.94x | ** 5.25x** | ** 5.74x** | ** 2.93x** |
131- | 30 | 2.67x | ** 5.59x** | ** 6.21x** | ** 3.07x** |
119+ | ` --n-cpu-moe ` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
120+ | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
121+ | 0 _ (baseline)_ | 16,822 | 14,602 | 1.15x | 11,173 | 10,843 | ** 1.03x** | 11,274 | 10,628 | ** 1.06x** | 161.4 | 206.7 | 0.78x |
122+ | 8 | 15,724 | 11,874 | 1.32x | 7,063 | 1,459 | ** 4.84x** | 6,500 | 1,459 | ** 4.46x** | 80.2 | 32.7 | ** 2.45x** |
123+ | 16 | 14,128 | 9,122 | 1.55x | 4,183 | 833 | ** 5.02x** | 4,888 | 854 | ** 5.72x** | 54.5 | 21.9 | ** 2.49x** |
124+ | 24 | 12,346 | 6,368 | 1.94x | 3,500 | 667 | ** 5.25x** | 3,958 | 689 | ** 5.74x** | 49.1 | 16.7 | ** 2.93x** |
125+ | 30 _ (` --cpu-moe ` )_ | 11,038 | 4,134 | 2.67x | 3,035 | 543 | ** 5.59x** | 3,072 | 495 | ** 6.21x** | 39.7 | 12.9 | ** 3.07x** |
132126
133127### Qwen 3.5 35B-A3B (UD-IQ4_XS, 48 MoE layers)
134128
135- | ` --n-cpu-moe ` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
136- | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
137- | 0 _ (baseline)_ | 19,862 | 9,538 | 9,405 | 160.0 | 17,522 | 8,149 | 8,073 | 228.4 |
138- | 12 | 18,148 | 6,755 | 6,648 | 75.4 | 13,282 | 988 | 954 | 27.5 |
139- | 24 | 15,414 | 4,412 | 5,259 | 52.3 | 9,010 | 498 | 484 | 15.8 |
140- | 36 | 12,684 | 3,772 | 4,223 | 50.7 | 4,738 | 523 | 517 | 11.3 |
141- | 48 _ (` --cpu-moe ` )_ | 11,606 | 3,917 | 3,709 | 38.6 | 3,314 | 477 | 457 | 10.2 |
142-
143- | ` --n-cpu-moe ` | VRAM | pp4096 | pp8192 | tg128 |
144- | ---| ---:| ---:| ---:| ---:|
145- | 0 | 1.13x | ** 1.17x** | ** 1.16x** | 0.70x |
146- | 12 | 1.37x | ** 6.84x** | ** 6.97x** | ** 2.74x** |
147- | 24 | 1.71x | ** 8.85x** | ** 10.86x** | ** 3.31x** |
148- | 36 | 2.68x | ** 7.21x** | ** 8.17x** | ** 4.50x** |
149- | 48 | 3.50x | ** 8.21x** | ** 8.11x** | ** 3.77x** |
129+ | ` --n-cpu-moe ` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
130+ | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
131+ | 0 _ (baseline)_ | 19,862 | 17,522 | 1.13x | 9,538 | 8,149 | ** 1.17x** | 9,405 | 8,073 | ** 1.16x** | 160.0 | 228.4 | 0.70x |
132+ | 12 | 18,148 | 13,282 | 1.37x | 6,755 | 988 | ** 6.84x** | 6,648 | 954 | ** 6.97x** | 75.4 | 27.5 | ** 2.74x** |
133+ | 24 | 15,414 | 9,010 | 1.71x | 4,412 | 498 | ** 8.85x** | 5,259 | 484 | ** 10.86x** | 52.3 | 15.8 | ** 3.31x** |
134+ | 36 | 12,684 | 4,738 | 2.68x | 3,772 | 523 | ** 7.21x** | 4,223 | 517 | ** 8.17x** | 50.7 | 11.3 | ** 4.50x** |
135+ | 48 _ (` --cpu-moe ` )_ | 11,606 | 3,314 | 3.50x | 3,917 | 477 | ** 8.21x** | 3,709 | 457 | ** 8.11x** | 38.6 | 10.2 | ** 3.77x** |
150136
151137### GPT-OSS 20B (Q8_0 / MXFP4, 24 MoE layers)
152138
153- | ` --n-cpu-moe ` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
154- | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
155- | 0 _ (baseline)_ | 13,186 | 13,964 | 12,925 | 212.8 | 12,204 | 17,856 | 17,642 | 344.2 |
156- | 6 | 11,560 | 8,975 | 7,617 | 85.8 | 9,812 | 1,747 | 1,666 | 32.2 |
157- | 12 | 9,378 | 6,470 | 6,394 | 51.7 | 7,386 | 1,176 | 1,188 | 18.3 |
158- | 18 | 7,192 | 4,315 | 4,393 | 30.7 | 4,962 | 807 | 751 | 12.1 |
159- | 24 _ (` --cpu-moe ` )_ | 4,762 | 4,277 | 3,798 | 27.7 | 2,536 | 568 | 548 | 9.4 |
160-
161- | ` --n-cpu-moe ` | VRAM | pp4096 | pp8192 | tg128 |
162- | ---| ---:| ---:| ---:| ---:|
163- | 0 | 1.08x | 0.78x | 0.73x | 0.62x |
164- | 6 | 1.18x | ** 5.14x** | ** 4.57x** | ** 2.67x** |
165- | 12 | 1.27x | ** 5.50x** | ** 5.38x** | ** 2.83x** |
166- | 18 | 1.45x | ** 5.35x** | ** 5.85x** | ** 2.54x** |
167- | 24 | 1.88x | ** 7.53x** | ** 6.93x** | ** 2.95x** |
139+ | ` --n-cpu-moe ` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
140+ | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
141+ | 0 _ (baseline)_ | 13,186 | 12,204 | 1.08x | 13,964 | 17,856 | 0.78x | 12,925 | 17,642 | 0.73x | 212.8 | 344.2 | 0.62x |
142+ | 6 | 11,560 | 9,812 | 1.18x | 8,975 | 1,747 | ** 5.14x** | 7,617 | 1,666 | ** 4.57x** | 85.8 | 32.2 | ** 2.67x** |
143+ | 12 | 9,378 | 7,386 | 1.27x | 6,470 | 1,176 | ** 5.50x** | 6,394 | 1,188 | ** 5.38x** | 51.7 | 18.3 | ** 2.83x** |
144+ | 18 | 7,192 | 4,962 | 1.45x | 4,315 | 807 | ** 5.35x** | 4,393 | 751 | ** 5.85x** | 30.7 | 12.1 | ** 2.54x** |
145+ | 24 _ (` --cpu-moe ` )_ | 4,762 | 2,536 | 1.88x | 4,277 | 568 | ** 7.53x** | 3,798 | 548 | ** 6.93x** | 27.7 | 9.4 | ** 2.95x** |
168146
169147> GPT-OSS is the one model still behind on the resident path at these prompt
170148> lengths (0.73–0.78x). Moving to 4,096/8,192-token prompts is what exposed it:
@@ -183,17 +161,11 @@ DeepSeek V4 does not use the host-MoE seam: it hands its offloaded layers to
183161` ggml_backend_sched ` , which streams the weights back to the GPU for a
184162prefill-sized batch exactly the way llama.cpp does.
185163
186- | ` --n-cpu-moe ` | TS VRAM (MiB) | TS pp4096 | TS pp8192 | TS tg128 | llama VRAM (MiB) | llama pp4096 | llama pp8192 | llama tg128 |
187- | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
188- | 0 _ (baseline, both GPUs)_ | 169,132 | 3,448 | 4,387 | 51.1 | 155,608 | 2,398 | 2,232 | 49.6 |
189- | 12 | 131,818 | 392 | 428 | 10.3 | 117,150 | 126 | 124 | 13.7 |
190- | 24 | 79,742 | 218 | 236 | 5.3 | 78,954 | 64 | 63 | 7.2 |
191-
192- | ` --n-cpu-moe ` | VRAM | pp4096 | pp8192 | tg128 |
193- | ---| ---:| ---:| ---:| ---:|
194- | 0 | 1.09x | ** 1.44x** | ** 1.97x** | ** 1.03x** |
195- | 12 | 1.13x | ** 3.11x** | ** 3.46x** | 0.75x |
196- | 24 | 1.01x | ** 3.42x** | ** 3.72x** | 0.74x |
164+ | ` --n-cpu-moe ` | TS VRAM (MiB) | llama VRAM (MiB) | ratio | TS pp4096 | llama pp4096 | ratio | TS pp8192 | llama pp8192 | ratio | TS tg128 | llama tg128 | ratio |
165+ | ---| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:| ---:|
166+ | 0 _ (baseline, both GPUs)_ | 169,132 | 155,608 | 1.09x | 3,448 | 2,398 | ** 1.44x** | 4,387 | 2,232 | ** 1.97x** | 51.1 | 49.6 | ** 1.03x** |
167+ | 12 | 131,818 | 117,150 | 1.13x | 392 | 126 | ** 3.11x** | 428 | 124 | ** 3.46x** | 10.3 | 13.7 | 0.75x |
168+ | 24 | 79,742 | 78,954 | 1.01x | 218 | 64 | ** 3.42x** | 236 | 63 | ** 3.72x** | 5.3 | 7.2 | 0.74x |
197169
198170DeepSeek V4 remains the one model where llama.cpp's offloaded * decode* is ahead
199171(0.74–0.75x): both engines run that matmul through the same ggml CPU backend, and
0 commit comments