Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Dialogue that can listen, speak, and appear.

Haoyu Zhang*1 Zhipeng Li2 Xiaoying Tang1 Tianshu Yu†1 Yiwen Guo†3
1 The Chinese University of Hong Kong, Shenzhen 2 LIGHTSPEED 3 Independent Researcher
Reference portraits and generated expressive response frames
Dialogue-native visual responses. Given a multimodal query, reference image, and reference audio, Ex-Omni-2D generates coordinated text, personalized speech, and synchronized video.
How it works

Method Overview

A dialogue model plans visual intent and produces speech units; speech and video generators realize one synchronized response.

Ex-Omni-2D framework overview
Ex-Omni-2D connects multimodal understanding, visual planning, native speech generation, and reference-conditioned video synthesis.
01

Multimodal Character Grounding

Speech, text, and reference vision provide conversational context, identity, appearance, and voice conditions.

02

Visual Response Planning

The LLM emits a structured VTP for first frame, scene, emotion, movement style, and motion details.

03

Native Speech Interface

Sixteen codebooks generate personalized speech and serve as a compact, frame-aligned video condition.

04

Teacher & Streaming Student

A full-sequence Teacher prioritizes quality; its Prefix-Streaming Student enables efficient incremental deployment.

Generated responses

Qualitative Results

Ex-Omni-2D preserves reference appearance while realizing dialogue-aware expression and speech-aligned motion.

Full-sequence quality

Video Demos

Prefix Streaming

Efficient incremental generation

Streaming Video Demos

Visual comparison without and with Prefix Streaming

Late-chunk appearance

Prefix Streaming reduces cumulative deformation while retaining reference identity.

Long-horizon consistency plots

16-chunk consistency

Prefix is stronger from chunks 9–16 and reduces last-to-first consistency error by 21.4%.

Reference

Citation

If you find this work useful, please cite our paper and related prior work.

BibTeX
@article{zhang2026exomni2d,
  title  = {Ex-Omni-2D: Expressive Omni-Modal Dialogue
            Models with Native Visual Presence},
  author = {Zhang, Haoyu and Li, Zhipeng and Tang, Xiaoying
            and Yu, Tianshu and Guo, Yiwen},
  journal = {arXiv preprint arXiv:2608.10720},
  year   = {2026}
}

@article{zhang2026ex,
  title   = {Ex-Omni: Enabling 3D Facial Animation Generation
             for Omni-modal Large Language Models},
  author  = {Zhang, Haoyu and Li, Zhipeng and Guo, Yiwen
             and Yu, Tianshu},
  journal = {arXiv preprint arXiv:2602.07106},
  year    = {2026}
}