Research Arena is DefChat's automated research evaluation protocol. Each episode selects two AI research papers relevant to voice identity, agent trust, and episodic cognition, then runs five independent trials through three specialist judges — Voice & Identity, Trust & Provenance, and Episodic Cognition — and computes a weighted verdict using a calibrated reducer. The full decision artefact is available for download below.
Three independent judge personas scored each paper on Novelty, Evidence, and Impact (0–10). Total is the normalised composite (0–1). Scores shown as Paper A / Paper B. Scores shown from trial 1. Full trial data available in the Proof-of-Decision JSON.
| Judge | Novelty | Evidence | Impact | Total (A / B) | Confidence | Reliability |
|---|---|---|---|---|---|---|
|
Voice_Identity
Voice & Identity · claude-haiku
|
4.50
/
3.00
|
6.00
/
4.00
|
5.50
/
2.00
|
0.53
/
0.30
|
0.75 | Medium |
|
A: This paper applies existing LLM + rule-based validation techniques to a socially valuable but narrow domain (maternal health enrollment) without advancing voice identity, speaker verification, consent architecture, or protection against voice fraud—the core concerns of voice-as-identity research.
B: This work addresses audio-visual source localization through self-supervised learning, which is orthogonal to voice identity, speaker verification, consent architecture, or voice fraud detection—the core domains of voice identity research.
|
||||||
|
Trust_Provenance
Trust & Provenance · claude-sonnet
|
4.00
/
2.00
|
5.50
/
3.00
|
3.00
/
1.00
|
0.42
/
0.20
|
0.78 | High |
|
A: FormBharo addresses a real deployment challenge but contributes little to trust infrastructure or provenance mechanisms, as its rule-based validation and benchmark are domain-specific engineering rather than advances in auditability, agent accountability, or inspectable AI decision-making.
B: This audio-visual source localization paper has no meaningful connection to AI trust infrastructure, agent accountability, or provenance chains, making it essentially irrelevant to the evaluation domain.
|
||||||
|
Episodic_Cognition
Episodic Cognition · openai-4o
|
5.00
/
6.00
|
6.00
/
7.00
|
5.00
/
5.00
|
0.53
/
0.60
|
0.80 | High |
|
A: The paper presents an incremental improvement by integrating rule-based validation with LLMs for voice agents, validated on a new benchmark, but its impact is limited to a specific application domain.
B: The paper introduces a novel approach to audio-visual learning but lacks direct relevance to episodic memory or long-horizon reasoning.
|
||||||
5 independent evaluations were run. Each trial ran the full judge panel independently.
| Trial | Winner | Margin | Confidence | Agreement |
|---|---|---|---|---|
| 1 | Paper A | 0.152 | 0.243 | ✓ Majority |
| 2 | Paper A | 0.118 | 0.189 | ✓ Majority |
| 3 | Paper A | 0.081 | 0.162 | ✓ Majority |
| 4 | Paper A | 0.127 | 0.203 | ✓ Majority |
| 5 | Paper A | 0.133 | 0.212 | ✓ Majority |
This decision was computed using normalised scores, confidence weighting, and calibration-adjusted reliability across three independent judges. The reducer applies a weighted aggregation before producing a final margin and confidence estimate.
↓ Download Proof-of-Decision (JSON)