Liva-Party-Bench
Liva-Party-Bench
Evaluating Full-Duplex Models in Group Conversation Settings
We believe the future of conversational voice AI goes beyond 1-on-1. AI will join meetings, hang out with a group of friends, and assist families at home. To get there, models need to handle group dynamics like knowing when to stay quiet, recognizing who's being addressed, and responding at the right moment. Liva-Party-Bench evaluates whether full-duplex models can navigate these multi-party settings, built from real 3+ speaker conversations across English, Hindi, and Arabic.
Last updated May 28, 2026
Group Social Awareness
Can the model detect private/sensitive topics among other participants and not barge in? Lower TOR is better; higher ratings are better. NSA = not socially acceptable and SA = socially acceptable for model interjection.
Group Address Awareness
Can the model recognize when it is being addressed vs. when others are talking to each other? Higher TOR/rating is better; lower latency is better.
Methodology
Motivation. AI assistants and companions are often present during multi-party conversations, especially in shared spaces like homes, yet no prior benchmark evaluates full-duplex model performance in these settings. Liva-Party-Bench introduces two group conversational events that assess whether models can navigate the social dynamics of 3+ speaker conversations.
Source data. All samples are drawn from Yapdo, Liva AI's ~50,000-hour multilingual conversational speech corpus. Group samples come from recordings with 3+ speakers who are friends/acquaintances, natively recorded with separate speaker channels. 200 samples per language (English, Hindi, Arabic) per event.
Group social awareness. We identified 60-second windows where at least two speakers were actively conversing (10%+ speech activity each, with back-and-forth turn-taking) while a third speaker remained quiet (<5% activity). We ensured each clip had at least one 1s+ pause to give the model an opportunity to interject. Clips were transcribed with ElevenLabs Scribe v2 with speaker diarization and classified by an LLM as "socially acceptable" (SA) or "not socially acceptable" (NSA) for interjection.
Group address awareness. Using the same sliding window, we identified moments where two active speakers converse while a third speaker (C) remains quiet, then C enters the conversation near the end - within 5 seconds of the last utterance. An LLM verified that the active speakers were talking to each other (not C), that C's entry was explicitly prompted by an active speaker addressing C directly, and that there was a clear shift in addressee.
Context construction. Conversational context was transcribed with ElevenLabs Scribe v2 and summarized by an LLM from a third-party listener's perspective. For address awareness, the model's name was included since speakers often address each other by name. Context was provided as text.
Metrics. TOR NSA is the takeover rate in not-socially-acceptable scenarios (lower = model correctly stays quiet). LLM-judge ratings (0–5) evaluate whether the model chose the right moment and said appropriate things. Model latency is the time to respond when addressed (seconds).
Models evaluated. GPT-Realtime, Gemini 2.5 Flash Native Audio, Amazon Nova Sonic, and Grok. Results averaged over 200 samples per language.