ACM SIGGRAPH MIG 2026 · Charleston, SC

Constructing Dynamic Conversations with Neural Back-channel Generation and Voice-line Scheduling

A
B

The Neural Scheduler places every clip in time. Replies can overlap, pauses vary in length, and the back-channel lands where line 3 was split.

Schematic, not measured data. Speaker A is amber and speaker B is blue; numbers give the script order.

Overview video

Abstract

We present a novel approach to the problem of combining two voice-line channels of dyadic conversational audio into a unified conversation with controllable back-channeling and turn-taking dynamics.

In voice performances for the game industry, the two tracks of voice-lines are typically recorded or synthesized in isolation. Our method takes two such unsynchronized tracks and uses a neural scheduler to sequence them into a single dynamic conversation, modifying both tracks by inserting or removing pauses and introducing voice overlap as appropriate. We further propose a back-channel predictor that is able to predict opportunity for back-channels for a given conversational turn's audio and content, which the scheduler can then incorporate.

We evaluate our scheduler and predictor independently and in combination by: various quantitative comparison; a qualitative and perceptual comparison to prior art and naive baselines; a compelling set of results showing animator control over the diversity, extent and nature of back-channeling and turn dynamics.

How it works

Pipeline diagram. Input voice clips from speakers A and B pass through the Back-channel Predictor, which adds a generated back-channel clip, and then the Neural Scheduler, which arranges all clips on a two-channel timeline. Four insets show the animator controls: clip splitting, the back-channel probability threshold, the back-channel vocabulary, and timing constraints.
The pipeline takes voice lines in script order. Circled numbers mark the four controls an animator can set.
Component 1

Back-channel Predictor

A classifier reads each clip's text and audio and decides whether it invites a listener response, and whether that response should be a quiet acknowledgment (“uh-huh”) or a stronger reaction (“wow, really?”). The back-channel itself comes from a vocabulary the animator supplies.

Component 2

Neural Scheduler

A full-duplex language model that is given the content of every line and learns only its timing. We treat scheduling as next-token prediction over the two speakers' channels.

Inference

Search with controls

A three-pass search uses the scheduler to score candidate start times for each clip. Its settings are the animator's controls, so behavior changes without retraining.

Controlling the conversation

Each group below is one dialogue rendered under different settings. The faces are animated from the scheduled audio.

Turn timing

Default scheduling. The scheduler picks each gap freely.
Tight scheduling. Onset constraints are tightened, so replies come sooner and overlap more.

Back-channels

No back-channels. The listener stays silent.
With back-channels. Brief acknowledgments are inserted where the predictor finds an opening.
Salient back-channels. The candidate pool is swapped for longer, more noticeable responses.

Results

Back-channels on Full-Duplex Bench

TOR is the share of back-channels that take over the speaker's turn, Freq is back-channels per second, and JSD is the distance from the timing of human back-channels.
ModelTOR ↓Freq ↑JSD ↓
Freeze-Omni0.6360.0010.997
Gemini0.0910.0120.896
Moshi1.0000.0010.957
dGSLM0.6910.0150.934
PersonaPlex0.3270.0250.649
Ours0.0000.0620.644

Our TOR is zero by construction, because the system only inserts back-channels and never hands the listener the floor.

Which sounds more natural?

Share of forced-choice trials in which 20 listeners preferred our output. Whiskers are 95% confidence intervals and the dashed line is chance. The casual-chat result is not significant (p = 0.617): with simple, sequential turns the timing difference is hard to hear.

Voice activity timeline over 20 seconds comparing a fixed-pause baseline (top) with our scheduler (bottom) on the same dialogue. Highlighted regions show that our schedule contains overlapping speech and gaps of varied length, whereas the baseline's gaps are uniform.
One dialogue scheduled two ways. The fixed-pause baseline (top, blue) spaces every turn evenly, while our scheduler (bottom, orange) produces overlaps and gaps of varied length.
Two voice activity timelines, Examples A and B, each showing a speaker's turns together with the back-channels produced by our model and by PersonaPlex. Our back-channels are short and fall within the speaker's turns, while PersonaPlex's are long or come after the speaker has finished.
Back-channels from our model (orange) and PersonaPlex (green) against the speaker's turns (blue). PersonaPlex either waits until the speaker has finished (A) or takes over the floor (B).

BibTeX

@inproceedings{pan2026voiceline,
  author    = {Pan, Yifang and Landreth, Chris and Singh, Karan},
  title     = {Constructing Dynamic Conversations with Neural Back-channel
               Generation and Voice-line Scheduling},
  booktitle = {The 19th ACM SIGGRAPH Conference on Motion, Interaction,
               and Games (MIG '26)},
  year      = {2026},
  address   = {Charleston, SC, USA},
  publisher = {ACM},
  isbn      = {979-8-4007-2824-2},
  doi       = {10.1145/3828647.3857292}
}