VLDB 2026 Research / reviewers in the wild / expert
Zixi Tuo
dblp:342/8032
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0005-5402-0458ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 90% Language models and text generation · 6% Deep learning architectures and training · 4% | |
| Computer graphics and multimedia
4 papers |
Visual content generation and editing · 73% Image and video processing · 17% Multimedia analysis and retrieval · 10% |
Topics — the 7 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.4 | 3 | 2025 | DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation · Int. J. Comput. Vis. 2025 MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text · ACM Multimedia 2023 |
Visual content generation and editing › video generation
text-to-video generation |
1.3 | 2 | 2023 | MovieFactory: Automatic Movie Creation from Text using Large Generative Models for Language and Images · ACM Multimedia 2023 MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text · ACM Multimedia 2023 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.9 | 1 | 2025 | Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation · Int. J. Comput. Vis. 2025 |
Visual content generation and editing › multimodal content generation
story visualization |
0.9 | 1 | 2025 | DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Image and video processing › super-resolution
video super-resolution |
0.7 | 1 | 2023 | Learning Data-Driven Vector-Quantized Degradation Model for Animation Video Super-Resolution · ICCV 2023 |
Multimedia analysis and retrieval
audio retrieval |
0.4 | 2 | 2023 | MovieFactory: Automatic Movie Creation from Text using Large Generative Models for Language and Images · ACM Multimedia 2023 MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text · ACM Multimedia 2023 |
Natural language and speech › Language models and text generation › text generation
LLM-guided generation |
0.3 | 1 | 2025 | DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.0masked mutual self-attention · 1.7masked mutual cross-attention · 1.7vector quantization · 1.3text-to-image generation · 1.3data enhancement · 1.3codebook learning · 1.3multimodal anchors · 0.9multimodal anchor · 0.9attention mechanism · 0.9temporal learning · 0.7large language model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
Wenjing Wang 0001, Huan Yang 0005, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, Jiaying Liu 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent DiffusionabstractStory visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results, existing methods still struggle to create a coherent sequence of subject-consistent frames based solely on a story. To this end, we propose DreamStory, an automatic open-domain story visualization framework by leveraging the LLMs and a novel multi-subject consistent diffusion model. DreamStory consists of (1) an LLM acting as a story director and (2) an innovative Multi-Subject consistent Diffusion model (MSD) for generating consistent multi-subject across the images. First, DreamStory employs the LLM to generate descriptive prompts for subjects and scenes aligned with the story, annotating each scene's subjects for subsequent subject-consistent generation. Second, DreamStory utilizes these detailed subject descriptions to create portraits of the subjects, with these portraits and their corresponding textual information serving as multimodal anchors (guidance). Finally, the MSD uses these multimodal anchors to generate story scenes with consistent multi-subject. Specifically, the MSD includes Masked Mutual Self-Attention (MMSA) and Masked Mutual Cross-Attention (MMCA) modules. MMSA module ensures detailed appearance consistency with reference images, while MMCA captures key attributes of subjects from their reference text to ensure semantic consistency. Both modules employ masking mechanisms to restrict each scene's subjects to referencing the multimodal information of the corresponding subject, effectively preventing blending between multiple subjects. To validate our approach and promote progress in story visualization, we established a benchmark, DS-500, which can assess the overall performance of the story visualization framework, subject-identification accuracy, and the consistency of the generation model. Extensive experiments validate the effectiveness of DreamStory in both subjective and objective evaluations. Huiguo He, Huan Yang 0005, Zixi Tuo, Qiuyue Wang, Wenhao Huang 0001, Hongyang Chao, Jian Yin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Learning Data-Driven Vector-Quantized Degradation Model for Animation Video Super-ResolutionabstractExisting real-world video super-resolution (VSR) methods focus on designing a general degradation pipeline for open-domain videos while ignoring data intrinsic characteristics which strongly limit their performance when applying to some specific domains (e.g., animation videos). In this paper, we thoroughly explore the characteristics of animation videos and leverage the rich priors in real-world animation data for a more practical animation VSR model. In particular, we propose a multi-scale Vector-Quantized Degradation model for animation video Super-Resolution (VQD-SR) to decompose the local details from global structures and transfer the degradation priors in real-world animation videos to a learned vector-quantized codebook for degradation modeling. A rich-content Real Animation Low-quality (RAL) video dataset is collected for extracting the priors. We further propose a data enhancement strategy for high-resolution (HR) training videos based on our observation that existing HR videos are mostly collected from the Web which contains conspicuous compression artifacts. The proposed strategy is valid to lift the upper bound of animation VSR performance, regardless of the specific VSR model. Experimental results demonstrate the superiority of the proposed VQD-SR over state-of-the-art methods, through extensive quantitative and qualitative evaluations of the latest animation video super-resolution benchmark. The code and pre-trained models can be downloaded at https://github.com/researchmm/VQD-SR. Zixi Tuo, Huan Yang 0005, Jianlong Fu, Yujie Dun, Xueming Qian |
ICCV | 1 |
| 2023 | MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from TextabstractVideos for mobile devices become the most popular access to share and acquire information recently. For the convenience of users' creation, in this paper, we present a system, namely MobileVidFactory, to automatically generate vertical mobile videos where users only need to give simple texts mainly. Our system consists of two parts: basic and customized generation. In the basic generation, we utilize the pretrained image diffusion model, and adapt it to a high-quality open-domain vertical video generator. As for the audio, by retrieving from our big database, our system matches a suitable background sound for the video. Additionally to produce customized content, our system allows users to add specified screen texts for enriching visual expression, and specify texts for automatic reading with optional voices as they like. Junchen Zhu, Huan Yang 0005, Wenjing Wang 0001, Huiguo He, Zixi Tuo, Wen-Huang Cheng, Lianli Gao, Jingkuan Song, Jianlong Fu, Jiebo Luo 0001 |
ACM Multimedia | 5 |
| 2023 | MovieFactory: Automatic Movie Creation from Text using Large Generative Models for Language and ImagesabstractIn this paper, we present MovieFactory, a powerful framework to generate cinematic-picture (3072x1280), film-style (multi-scene), and multi-modality (sounding) movies on the demand of natural languages. As the first fully automated movie generation model to the best of our knowledge, our approach empowers users to create captivating movies with smooth transitions using simple text inputs, surpassing existing methods that produce soundless videos limited to a single scene of modest quality. To facilitate this distinctive functionality, we leverage ChatGPT to expand user-provided text into detailed sequential scripts for movie generation. Then we bring scripts to life visually and acoustically through vision generation and audio retrieval. To generate videos, we extend the capabilities of a pretrained text-to-image diffusion model through a two-stage process. Firstly, we employ spatial finetuning to bridge the gap between the pretrained image model and the new video dataset. Subsequently, we introduce temporal learning to capture object motion. In terms of audio, we leverage sophisticated retrieval models to select and align audio elements that correspond to the plot and visual content of the movie. Junchen Zhu, Huan Yang 0005, Huiguo He, Wenjing Wang 0001, Zixi Tuo, Wen-Huang Cheng, Lianli Gao, Jingkuan Song, Jianlong Fu |
ACM Multimedia | 5 |