Research Arena

Episode Verdict
episodef40fb4e3-7715-496d-9ecb-019728bfd66f
judge_v2.0
reducer_v1.1
ts2026-08-08T08:59:06.207918+00:00

Research Arena is DefChat's automated research evaluation protocol. Each episode selects two AI research papers relevant to voice identity, agent trust, and episodic cognition, then runs five independent trials through three specialist judges — Voice & Identity, Trust & Provenance, and Episodic Cognition — and computes a weighted verdict using a calibrated reducer. The full decision artefact is available for download below.

Judges: Voice & Identity · Trust & Provenance · Episodic Cognition Trials per episode: 5 Reducer: Confidence-weighted, disagreement-dampened

Evaluated Papers

Paper A
FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India
Aman Dalmia, Sanskriti Midha, Jigar Doshi
Published Aug 6, 2026
# Summary The researchers built FormBharo, a voice agent that uses Large Language Models paired with rule-based validation to help illiterate Hindi-speaking mothers in India enroll in maternal and child health programs over phone calls, which is now being piloted with an NGO called ARMMAN. They created FormVoiceAgentBench, a benchmark with 3,760 multi-turn conversation tests, and found that form completion accuracy drops significantly (~41 points) with real speech transcripts versus perfect ones, but rule-based controls help smaller models match larger ones, and end-to-end evaluation is necessary because component performance doesn't predict overall success.
Paper B
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
Han Hu, Dongheng Lin, Yuqi Hou
Published Aug 6, 2026
# Summary The researchers discovered that self-supervised audio-visual learning models naturally focus on the most salient sound source (selective convergence), which they leveraged in a two-stage framework to progressively identify dominant and then remaining sound sources without manual annotations. They achieved competitive performance on dual-source localization benchmarks and identified evaluation inconsistencies in existing benchmarks, proposing pixel-level segmentation masks for more accurate spatial evaluation.

Judge Evaluation

Three independent judge personas scored each paper on Novelty, Evidence, and Impact (0–10). Total is the normalised composite (0–1). Scores shown as Paper A / Paper B. Scores shown from trial 1. Full trial data available in the Proof-of-Decision JSON.

Judge Novelty Evidence Impact Total (A / B) Confidence Reliability
Voice_Identity
Voice & Identity · claude-haiku
4.50 / 3.00
6.00 / 4.00
5.50 / 2.00
0.53 / 0.30
0.75 Medium
A: This paper applies existing LLM + rule-based validation techniques to a socially valuable but narrow domain (maternal health enrollment) without advancing voice identity, speaker verification, consent architecture, or protection against voice fraud—the core concerns of voice-as-identity research.
B: This work addresses audio-visual source localization through self-supervised learning, which is orthogonal to voice identity, speaker verification, consent architecture, or voice fraud detection—the core domains of voice identity research.
Trust_Provenance
Trust & Provenance · claude-sonnet
4.00 / 2.00
5.50 / 3.00
3.00 / 1.00
0.42 / 0.20
0.78 High
A: FormBharo addresses a real deployment challenge but contributes little to trust infrastructure or provenance mechanisms, as its rule-based validation and benchmark are domain-specific engineering rather than advances in auditability, agent accountability, or inspectable AI decision-making.
B: This audio-visual source localization paper has no meaningful connection to AI trust infrastructure, agent accountability, or provenance chains, making it essentially irrelevant to the evaluation domain.
Episodic_Cognition
Episodic Cognition · openai-4o
5.00 / 6.00
6.00 / 7.00
5.00 / 5.00
0.53 / 0.60
0.80 High
A: The paper presents an incremental improvement by integrating rule-based validation with LLMs for voice agents, validated on a new benchmark, but its impact is limited to a specific application domain.
B: The paper introduces a novel approach to audio-visual learning but lacks direct relevance to episodic memory or long-horizon reasoning.

Trial Results

5 independent evaluations were run. Each trial ran the full judge panel independently.

Trial Winner Margin Confidence Agreement
1 Paper A 0.152 0.243 ✓ Majority
2 Paper A 0.118 0.189 ✓ Majority
3 Paper A 0.081 0.162 ✓ Majority
4 Paper A 0.127 0.203 ✓ Majority
5 Paper A 0.133 0.212 ✓ Majority

Verdict

Research Arena — Episode Verdict

Winner Paper A
Agreement 5 / 5 trials
Avg Margin 0.122
Meta Confidence 0.20 Weak

This decision was computed using normalised scores, confidence weighting, and calibration-adjusted reliability across three independent judges. The reducer applies a weighted aggregation before producing a final margin and confidence estimate.

↓ Download Proof-of-Decision (JSON)