-
Notifications
You must be signed in to change notification settings - Fork 1.4k
Pull requests: flashinfer-ai/flashinfer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
perf(prefill): reduce with Tensor.max() instead of the Python builtin in plan()
op: attention
#5043
opened Sep 9, 2026 by
gf239
Loading…
5 tasks done
feat(mla): add CUDA graph plan update API
op: attention
#5041
opened Sep 8, 2026 by
saltyminty
Collaborator
Loading…
5 of 11 tasks
refactor(kda): stop 0.7 advertising unreleased KDA surface (mark prefill wrapper experimental, trim kda_kernels.__all__)
op: linear attention
KDA, mamba, GDN, etc. review filtering.
#5040
opened Sep 8, 2026 by
kahyunnam
Member
Loading…
10 tasks done
fix(attention): default skip_all_rows_active_check to true
op: attention
run-ci
#5039
opened Sep 8, 2026 by
saltyminty
Collaborator
Loading…
3 of 11 tasks
fix(decode): reserve max(K+V, st.o) smem for FA2 FP8+GQA decode
op: attention
run-ci
#5038
opened Sep 8, 2026 by
ir1ka
Contributor
Loading…
5 of 11 tasks
fix(kda): make recurrent_kda backend="auto" decode fall back to CuTe DSL
op: linear attention
KDA, mamba, GDN, etc. review filtering.
run-ci
v0.6.19
#5037
opened Sep 8, 2026 by
kahyunnam
Member
Loading…
5 tasks done
test(kda): restore the frozen module-ident consistency check
op: linear attention
KDA, mamba, GDN, etc. review filtering.
run-ci
#5035
opened Sep 8, 2026 by
kahyunnam
Member
Loading…
5 tasks done
fix(sampling): make multi-CTA top-k renorm deterministic (fixed-order cross-CTA sum)
op: misc
norm, activation, sampling, RoPE, quantization, etc.
#5034
opened Sep 8, 2026 by
gilfordting
•
Draft
3 tasks done
Add TOPK=256 to the SM120 NVFP4 sparse-MLA dispatch envelope
op: attention
#5033
opened Sep 8, 2026 by
danielwoz
Loading…
feat(kda): add CuTe DSL small-BH KDA prefill backend for Blackwell
op: linear attention
KDA, mamba, GDN, etc. review filtering.
#5032
opened Sep 8, 2026 by
qiangxu1996
Loading…
5 of 11 tasks
fix(moe): Fix per-token quantization
fast-math INF scale when activation contains all-zero row
op: moe
#5031
opened Sep 8, 2026 by
xuantengh
Contributor
Loading…
4 of 5 tasks
fix(attention): initialize variable-length attention sum accumulator
op: attention
#5029
opened Sep 8, 2026 by
shuo-ouyang
Loading…
5 of 11 tasks
fix(cake_kda): fix C16 forward triangular matrix orientation
op: linear attention
KDA, mamba, GDN, etc. review filtering.
#5028
opened Sep 8, 2026 by
zheyang0825
Loading…
4 of 11 tasks
feat(comm): add allocation-stable head-chunk Ulysses primitives
op: comm
#5027
opened Sep 8, 2026 by
tiffany940107
Contributor
Loading…
2 of 11 tasks
fix(moe): keep per-token FP8 quant scales finite for tiny activations
op: moe
#5024
opened Sep 8, 2026 by
yilin-void
Loading…
6 tasks done
feat(comm): add BF16 PCIe IPC all-gather and reduce-scatter
op: comm
#5023
opened Sep 8, 2026 by
yilin-void
•
Draft
Support compact GLM NoPE FP8 rows and fix masked reads
op: attention
#5022
opened Sep 8, 2026 by
ormandj
Contributor
Loading…
1 of 11 tasks
feat(kda): support graph-safe CuTe DSL prefix checkpoints
op: linear attention
KDA, mamba, GDN, etc. review filtering.
run-ci
#5021
opened Sep 8, 2026 by
Observer007
Contributor
Loading…
5 of 11 tasks
fix(attention): pre-scale P by 2^8 before e4m3 quantization in SM120 prims FP8 prefill
op: attention
#5020
opened Sep 8, 2026 by
wqyg18
Loading…
5 tasks done
test(attention): exercise the NVFP4 split-KV gated path on hardware
op: attention
#5018
opened Sep 7, 2026 by
Champollion9012
Loading…
5 of 11 tasks
fix(ci): preserve estimate file compression during refresh
#5014
opened Sep 7, 2026 by
mottopanikeiku
Loading…
4 of 11 tasks
fix(activation): pick a vector width that divides hidden_size in act_and_mul (fixes d=3420)
op: misc
norm, activation, sampling, RoPE, quantization, etc.
#5013
opened Sep 7, 2026 by
Wint3rNight
Loading…
5 tasks done
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.