-
Notifications
You must be signed in to change notification settings - Fork 814
Pull requests: open-compass/opencompass
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Dataset] Align ARC-AGI-1/2 evaluation with ARC Prize protocol
#2563
opened Jul 21, 2026 by
Vicentvankor
Loading…
[Feature] Add Multi-Round inferencer in GenInferencer and add Multi-IF dataset
#2557
opened Jul 17, 2026 by
Myhs-phz
Collaborator
Loading…
[Fix] Support PromptBench package for prompt attack
#2555
opened Jul 16, 2026 by
ssiq
Collaborator
Loading…
fix: use last-match in generic LLM judge postprocess (verdict-injection defense)
#2534
opened Jul 14, 2026 by
AUTHENSOR
Loading…
3 tasks done
Fix wikitext ppl evaluation without references
#2532
opened Jul 13, 2026 by
ssiq
Collaborator
Loading…
feat(dataset): add CHHallu-Src v1 — Chinese history hallucination & source-attribution benchmark
#2525
opened Jul 11, 2026 by
lizhuojunx86
Loading…
fix(MedCalc_Bench): guard calid-69 ground_truth regex match (AttributeError)
#2523
opened Jul 9, 2026 by
WatchTree-19
Loading…
fix(evaluator): guard IndexError on truncated judge choice (JudgeEvaluator/RMBEvaluator)
#2522
opened Jul 9, 2026 by
WatchTree-19
Loading…
Add robust HumanEval postprocess for chat outputs
#2515
opened Jul 8, 2026 by
Ding-god
Loading…
3 of 6 tasks
Add Helium Market Resolution benchmark
#2507
opened Jul 3, 2026 by
connerlambden
Loading…
1 of 2 tasks
[Fix] Extract the final answer in GPQA simple-eval predictions
#2496
opened Jun 27, 2026 by
Hibbert133
Contributor
Loading…
4 of 6 tasks
Add optional juryeval integration for LLM-as-Judge metrics
#2465
opened May 31, 2026 by
py-ai-dev
Loading…
[Fix] Combine split eval results in default summarizer
#2451
opened May 15, 2026 by
yhzhu99
Contributor
Loading…
[Fix] Use pre-tokenized prompts in VLLMwithChatTemplate to avoid modifying model input
#2434
opened Apr 15, 2026 by
suhmily10
Loading…
1 of 2 tasks
feat: upgrade MiniMax default model to M3
#2418
opened Mar 20, 2026 by
octo-patch
Contributor
Loading…
3 tasks done
[Fix] CEval ModelScope load and HF generate for causal LMs
#2416
opened Mar 19, 2026 by
DeliWang
Loading…
6 tasks
Previous Next
ProTip!
Adding no:label will show everything without a label.