Q4 gemm trail1 #16

shalinib-ibm · 2025-09-12T09:01:50Z

Approach 1 to optimise Q4 GEMM

Make sure to read the contributing guidelines before submitting a PR

This patch improves GEMM for FP32 Data Type on PowerPC Implements GEMM on large blocks with configurable block size mc, nc, kc (default: 256, 256, 256). Packing Function optimized to access blocks as per memory layout. GEMM Optimized to work on larger blocks. Isolated Packing from GEMM Operations for better MMA utilization. Verified functionality and correctness uing llama-cli and stand alone test case (performs matmul and compares final mattrix C result with base). Minor code refactoring changes: Replace macro with inline function Code Indent made consistent with 4 spaces Performance Testing: Observed 50% ~ 70% improvement in Prompt Processing Speed mesured using llama-bench with Meta-Llama3-8B FP32 Model. Similar gains observed with Mistral-7b-Instruct-v0.3 Model. model Size Params Backend Threads Test Patch Base llama 8B all F32 29.92 GiB 8.03 B CPU 20 pp512 98.58 60.3 llama 8B all F32 29.92 GiB 8.03 B CPU 20 pp1024 95.88 57.36 llama 8B all F32 29.92 GiB 8.03 B CPU 20 pp2048 85.46 53.26 llama 8B all F32 29.92 GiB 8.03 B CPU 20 pp4096 68.66 45.78 llama 8B all F32 29.92 GiB 8.03 B CPU 20 pp6144 57.35 40.44 25 ~ 30% improvement in llama-batched-bench with Metla-Llama3-8B in Prompt Processing Speed for large prompts (256, 512, 1024, 2048, 4096)tokens with various batch sizes ( 1, 2, 4, 8, 16) Signed-off-by: Shalini Salomi Bodapati <[email protected]>

CUrently in q4 gemm, a single block is being processed in KERNEL_mxn functions. For KERNEL_8x8, 8 rows -> each row, 1 block is being processed. I tried to modify this KERNEL_8x8 to process 8 rows -> 32 blocks once. Signed-off-by: Shalini Salomi Bodapati <[email protected]>

This reverts commit 48658b3.

shalinib-ibm added 2 commits August 26, 2025 03:10

github-actions bot added the ggml label Sep 12, 2025

Revert "PowerPC: Sgemm Optimization"

23d36bf

This reverts commit 48658b3.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Q4 gemm trail1 #16

Q4 gemm trail1 #16

Uh oh!

shalinib-ibm commented Sep 12, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Q4 gemm trail1 #16

Are you sure you want to change the base?

Q4 gemm trail1 #16

Uh oh!

Conversation

shalinib-ibm commented Sep 12, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants