Skip to content

Pull requests: open-compass/opencompass

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Fix LiveCodeBench evaluation failures
#2562 opened Jul 20, 2026 by ssiq Collaborator Loading…
Add MedFailBench dataset adapter
#2560 opened Jul 18, 2026 by goktugozkanmd Loading…
Add AA-LCR dataset integration
#2558 opened Jul 17, 2026 by ssiq Collaborator Loading…
Fix TopkRetriever metadata collation
#2556 opened Jul 17, 2026 by hongleng Loading…
5 of 6 tasks
[Fix] Support PromptBench package for prompt attack
#2555 opened Jul 16, 2026 by ssiq Collaborator Loading…
Fix vLLM chat template BOS handling
#2554 opened Jul 16, 2026 by ssiq Collaborator Loading…
Add BBH cascade evaluation config
#2553 opened Jul 16, 2026 by ssiq Collaborator Loading…
Fix livecodebench memory limit
#2538 opened Jul 14, 2026 by ssiq Collaborator Loading…
Fix wikitext ppl evaluation without references
#2532 opened Jul 13, 2026 by ssiq Collaborator Loading…
Add robust HumanEval postprocess for chat outputs
#2515 opened Jul 8, 2026 by Ding-god Loading…
3 of 6 tasks
Add Helium Market Resolution benchmark
#2507 opened Jul 3, 2026 by connerlambden Loading…
1 of 2 tasks
[Fix] Extract the final answer in GPQA simple-eval predictions
#2496 opened Jun 27, 2026 by Hibbert133 Contributor Loading…
4 of 6 tasks
Add BGPT REFUTE benchmark
#2471 opened Jun 5, 2026 by connerlambden Loading…
[Fix] Combine split eval results in default summarizer
#2451 opened May 15, 2026 by yhzhu99 Contributor Loading…
feat: upgrade MiniMax default model to M3
#2418 opened Mar 20, 2026 by octo-patch Contributor Loading…
3 tasks done
[Fix] CEval ModelScope load and HF generate for causal LMs
#2416 opened Mar 19, 2026 by DeliWang Loading…
6 tasks
ProTip! Adding no:label will show everything without a label.