[CUDA] Support FP8 (E4M3) KV Cache for Group Query Attention #10463
Triggered via pull request
February 14, 2026 04:12
Status
Success
Total duration
1h 21m 26s
Artifacts
1
windows_cuda.yml
on: pull_request
Windows GPU CUDA CI Pipeline
44m 9s
Windows GPU CUDA CI Pipeline Test Job
27m 6s
Annotations
6 warnings
|
Windows GPU CUDA CI Pipeline:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1234
epilog offset from end of function exceeds 4095
|
|
Windows GPU CUDA CI Pipeline:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1227
epilog offset from end of function exceeds 4095
|
|
Windows GPU CUDA CI Pipeline:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1220
epilog offset from end of function exceeds 4095
|
|
Windows GPU CUDA CI Pipeline:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1213
epilog offset from end of function exceeds 4095
|
|
Windows GPU CUDA CI Pipeline:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1206
epilog offset from end of function exceeds 4095
|
|
Windows GPU CUDA CI Pipeline:
onnxruntime/core/mlas/lib/amd64/QgemmU8X8KernelAvx2.asm#L1199
epilog offset from end of function exceeds 4095
|
Artifacts
Produced during runtime
| Name | Size | Digest | |
|---|---|---|---|
|
build-artifacts
|
2 GB |
sha256:8f1934677753de310beb81a15df387b4e4cb31f84c0d24975865aab4d979c4dc
|
|