fix(cuda): MTP + sm_89 compatibility for GCC 12 host compiler - #150
Merged
TheTom merged 1 commit intoJun 4, 2026
Merged
Conversation
- Disable -compress-mode flag (incompatible with GCC 12 host compiler) - Guard FP8 __nv_fp8_e4m3 conversion behind __CUDA_ARCH__ >= 900 (FP8 PTX instructions require sm_90+, causing build failure on RTX 4060 Ti / sm_89) Allows llama-server to build with CUDA on Ada Lovelace GPUs while retaining MTP (Multi-Token Prediction) support from upstream.
Owner
|
Thanks. The |
KGardevoir
pushed a commit
to KGardevoir/llama-cpp-turboquant
that referenced
this pull request
Jun 16, 2026
…heTom#150) - Disable -compress-mode flag (incompatible with GCC 12 host compiler) - Guard FP8 __nv_fp8_e4m3 conversion behind __CUDA_ARCH__ >= 900 (FP8 PTX instructions require sm_90+, causing build failure on RTX 4060 Ti / sm_89) Allows llama-server to build with CUDA on Ada Lovelace GPUs while retaining MTP (Multi-Token Prediction) support from upstream. Co-authored-by: savasuyar <savasuyar@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
-compress-modeflag — incompatible with GCC 12 host compiler (passed to GCC instead of nvcc)__nv_fp8_e4m3conversion behind__CUDA_ARCH__ >= 900— FP8 PTX instructions require sm_90+, causing build failure on RTX 4060 Ti / Ada Lovelace (sm_89)Why
Without these fixes, building with
-DGGML_CUDA=ONfails on sm_89 GPUs (RTX 4060 Ti, 4070, 4080, 4090) when using GCC 12 as the CUDA host compiler:gcc-12: error: unrecognized command-line option '-compress-mode=...'ptxas: Feature 'cvt with .e4m3x2/.e5m2x2' requires .target sm_90 or higherThese changes allow llama-server to build successfully with CUDA support on Ada Lovelace GPUs while retaining full MTP (Multi-Token Prediction) support from upstream.