Agent Review Panel is grounded in 9 peer-reviewed papers on multi-agent debate and evaluation quality.
| Paper | Venue | Contribution to This Skill |
|---|---|---|
| ChatEval | ICLR 2024 | Multi-agent debate evaluation with diverse role prompts |
| AutoGen | — | Solver/aggregator multi-agent architecture |
| Du et al. | ICML 2024 | Cross-verification through iterative debate for factuality |
| MachineSoM | ACL 2024 | Private reflection, conformity tracking |
| DebateLLM | ICML 2024 | Agreement intensity modulation |
| DMAD | ICLR 2025 | Diverse reasoning strategies per persona |
| Talk Isn't Always Cheap | ICML 2025 | Failure mode analysis informing debate quality safeguards |
| CONSENSAGENT | ACL 2025 | Dynamic sycophancy intervention |
| Trust or Escalate | ICLR 2025 Oral | Judge confidence gating with selective escalation |
The skill adapts these mechanisms (demonstrated on reasoning benchmarks) to the practical domain of code and document review. Additional integrations: AI Trust Evaluation Framework (claim verification, epistemic labels), VoltAgent (130+ specialist agents). See ROADMAP.md for planned additions.