Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
NVIDIA
/
TensorRT-LLM
Public
Notifications
You must be signed in to change notification settings
Fork
2.8k
Star
14.7k
Code
Issues
606
Pull requests
890
Discussions
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Discussions
Actions
Projects
Security and quality
Insights
Actions: NVIDIA/TensorRT-LLM
Actions
All workflows
Workflows
.github/workflows/test-runner.yml
.github/workflows/test-runner.yml
Auto Assign PR to Author
Auto Assign PR to Author
Auto Full Pre-Merge Approval Label
Auto Full Pre-Merge Approval Label
auto-assign
auto-assign
Blossom-CI
Blossom-CI
Bot-Command
Bot-Command
Clean Up Stale Pull Requests
Clean Up Stale Pull Requests
Close inactive issues
Close inactive issues
CodeQL
CodeQL
CodeRabbit Semantic Conflict Review
CodeRabbit Semantic Conflict Review
Show more workflows...
Management
Caches
Deployments
auto-assign
auto-assign
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Show workflow options
Create status badge
Create status badge
Loading
Uh oh!
There was an error while loading.
Please reload this page
.
auto-assign.yml
will be ignored since log searching is not yet available
1,047 workflow run results
1,047 workflow run results
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
[Usage]: how to deploy Qwen2.5-VL-7B-Instruct with triton-server
auto-assign
#1902:
Issue
#9142
labeled by
lqq-feel
6s
6s
View workflow file
[Feature][AutoDeploy]: Factory sharding as default
auto-assign
#1901:
Issue
#9137
labeled by
greg-kwasniewski1
1s
1s
View workflow file
[Feature][AutoDeploy]: MIxed EP+TP sharding
auto-assign
#1900:
Issue
#9136
labeled by
greg-kwasniewski1
1s
1s
View workflow file
speculative decoding in gemma3
auto-assign
#1899:
Issue
#6067
labeled by
karljang
1s
1s
View workflow file
[Feature]: KVCacheManager dump/load support
auto-assign
#1898:
Issue
#6962
labeled by
karljang
1s
1s
View workflow file
[Feature]: Add support for AFD in Step3 paper
auto-assign
#1897:
Issue
#7253
labeled by
karljang
1s
1s
View workflow file
[Feature]: Add all models to the AD dashboard
auto-assign
#1896:
Issue
#7399
labeled by
karljang
7s
7s
View workflow file
Function tooling like in vllm
auto-assign
#1895:
Issue
#6154
labeled by
karljang
1s
1s
View workflow file
[Feature]: AutoDeploy: Enable fp8 kv cache for Nano-v3 fp8 model
auto-assign
#1894:
Issue
#9102
labeled by
suyoggupta
1s
1s
View workflow file
[Bug]: AutoModelForCausalLM.
auto-assign
#1893:
Issue
#8942
labeled by
karljang
1s
1s
View workflow file
[Feature]: Support for _ and - parsing of CLI arguments in build_and_run_ad
auto-assign
#1892:
Issue
#9100
labeled by
greg-kwasniewski1
1s
1s
View workflow file
[Feature]: Sharding support for latent MoE
auto-assign
#1891:
Issue
#9098
labeled by
greg-kwasniewski1
2s
2s
View workflow file
[Bug]: Output includes tokens from other prompts when host kv cache enabled
auto-assign
#1890:
Issue
#8813
labeled by
karljang
2s
2s
View workflow file
[Feature][AutoDeploy] AOT compile moe_align kernel
auto-assign
#1889:
Issue
#9082
labeled by
pcastonguay
8s
8s
View workflow file
[Feature]: Add an option to configure the fused MoE backend from config YAML
auto-assign
#1887:
Issue
#9096
labeled by
nzmora-nvidia
1s
1s
View workflow file
[Feature]: Add an option to configure the fused MoE backend from config YAML
auto-assign
#1888:
Issue
#9096
labeled by
nzmora-nvidia
1s
1s
View workflow file
[Bug]: Tensor parallelism on A40 results in CUDA illegal instruction error
auto-assign
#1886:
Issue
#9086
labeled by
AlessioNetti
1s
1s
View workflow file
Crash During Stress Testing with Release v0.19.0rc0
auto-assign
#1885:
Issue
#3736
labeled by
karljang
Skipped
Skipped
View workflow file
Incompatibility between dp_attention and enable_chunked_prefill causes service hang on exceeding max_num_tokens
auto-assign
#1884:
Issue
#5267
labeled by
karljang
1s
1s
View workflow file
[Feature][AutoDeploy] AOT compile moe_align kernel
auto-assign
#1883:
Issue
#9082
labeled by
suyoggupta
8s
8s
View workflow file
How to enable deepep(especially cuda graph in generation) in PD disaggrated serving
auto-assign
#1882:
Issue
#5869
labeled by
karljang
1s
1s
View workflow file
Context node crash when using PD Disaggregation
auto-assign
#1881:
Issue
#3937
labeled by
karljang
1s
1s
View workflow file
The program freezes/hangs when using kv_cache_config=KvCacheConfig(event_buffer_max_size), batched_logits_processor=MyBatchedLogitsProcessor.
auto-assign
#1880:
Issue
#3857
labeled by
karljang
1s
1s
View workflow file
[Bug]: function cbapi->getCuptiStatus() failed with error CUPTI_ERROR_MULTIPLE_SUBSCRIBERS_NOT_SUPPORTED
auto-assign
#1879:
Issue
#9073
labeled by
zejunchen-zejun
1s
1s
View workflow file
Support for Qwen2-VL 2B/4B? or Qwen2.5-VL
auto-assign
#1878:
Issue
#8927
labeled by
yechank-nvidia
1s
1s
View workflow file
Previous
1
2
3
4
5
6
7
…
41
42
Next
You can’t perform that action at this time.