Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification
Under review
Citation (IEEE format): M. N. Chowdhury, N. J. Beg, U. H. Khan, S. R. Raiyan, M. K. Hasan and H. Mahmud, "Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification," arXiv preprint arXiv:2603.28488, 2026.
arXiv PDF Code/DataAuthors: Masnun Nuha Chowdhury†, Nusrat Jahan Beg†, Umme Hunny Khan, Syed Rifat Raiyan‡, Md Kamrul Hasan, Hasan Mahmud.
denotes equal contribution; denotes corresponding author.
Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augmented generation (RAG) and multi-agent debate (MAD) address this, they are limited by one-pass retrieval and unstructured debate dynamics. We propose a courtroom-style multi-agent framework, PROClaim, that reformulates verification as a structured, adversarial deliberation. Our approach integrates specialized roles (e.g., Plaintiff, Defense, Judge) with Progressive RAG (P-RAG) to dynamically expand and refine the evidence pool during the debate. Furthermore, we employ evidence negotiation, self-reflection, and heterogeneous multi-judge aggregation to enforce calibration, robustness, and diversity. In zero-shot evaluations on the Check-COVID benchmark, PROClaim achieves 81.7% accuracy, outperforming standard multi-agent debate by 10.0 percentage points, with P-RAG driving the primary performance gains (+7.5 pp). We show that the majority of this improvement stems from P-RAG’s dynamic coupling of retrieval to the evolving debate, rather than from any single component in isolation; the remaining courtroom mechanisms serve to stabilize and de-bias this process, together providing a robust foundation for reliable claim verification.
