Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zeqi Xiao

dblp:344/1615 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 54% Deep learning architectures and training · 10% Segmentation and scene understanding · 10%
Computer graphics and multimedia
4 papers
Visual content generation and editing · 64% Computer animation and physical simulation · 36%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
4.252025
WorldMem: Long-term Consistent World Simulation with Memory · NeurIPS 2025
Trajectory attention for fine-grained video motion control · ICLR 2025
TokensGen: Harnessing Condensed Tokens for Long Video Generation · ICCV 2025
Machine learning › Generative modeling › diffusion model
latent diffusion model
1.722025
TokensGen: Harnessing Condensed Tokens for Long Video Generation · ICCV 2025
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space · CVPR 2025
Visual content generation and editing
video generation
1.722025
WorldMem: Long-term Consistent World Simulation with Memory · NeurIPS 2025
Trajectory attention for fine-grained video motion control · ICLR 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
1.622025
WorldMem: Long-term Consistent World Simulation with Memory · NeurIPS 2025
Video Diffusion Models are Training-free Motion Interpreter and Controller · NeurIPS 2024
Computer animation and physical simulation
motion control
1.622025
Trajectory attention for fine-grained video motion control · ICLR 2025
Video Diffusion Models are Training-free Motion Interpreter and Controller · NeurIPS 2024
Machine learning › Generative modeling › video generation
long video generation
0.912025
TokensGen: Harnessing Condensed Tokens for Long Video Generation · ICCV 2025
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.912025
Position-Guided Point Cloud Panoptic Segmentation Transformer · Int. J. Comput. Vis. 2025
Computer vision › Segmentation and scene understanding › panoptic segmentation
point cloud panoptic segmentation
0.912025
Position-Guided Point Cloud Panoptic Segmentation Transformer · Int. J. Comput. Vis. 2025
Machine learning › Deep learning architectures and training › equivariant neural network
shift equivariance
0.912025
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space · CVPR 2025
Machine learning › Deep learning architectures and training › attention mechanism › visual attention
trajectory attention
0.912025
Trajectory attention for fine-grained video motion control · ICLR 2025
Machine learning › Generative modeling
video generation
0.912025
TokensGen: Harnessing Condensed Tokens for Long Video Generation · ICCV 2025
Visual content generation and editing › camera control
camera motion control
0.912025
Trajectory attention for fine-grained video motion control · ICLR 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.812024
CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics · NeurIPS 2024
Robotics › Motion planning and robot control › motion planning › contact planning
contact motion planning
0.812024
Unified Human-Scene Interaction via Prompted Chain-of-Contacts · ICLR 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
0.812024
CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics · NeurIPS 2024
Computer vision › 3D vision › 3d scene understanding › object relation reasoning
human-object interaction
0.812024
CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics · NeurIPS 2024
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding › human-centric scene understanding
human-scene interaction
0.812024
Unified Human-Scene Interaction via Prompted Chain-of-Contacts · ICLR 2024
Computer vision › 3D vision
point cloud analysis
0.312025
Position-Guided Point Cloud Panoptic Segmentation Transformer · Int. J. Comput. Vis. 2025
Visual content generation and editing
image-to-image translation
0.312025
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space · CVPR 2025
Machine learning › Trustworthy machine learning
interpretability
0.212024
Video Diffusion Models are Training-free Motion Interpreter and Controller · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning
0.212024
Unified Human-Scene Interaction via Prompted Chain-of-Contacts · ICLR 2024

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.6trajectory attention · 1.7state-aware memory attention · 1.7shift-equivariant attention · 1.7memory bank · 1.7equivariance loss · 1.7anti-aliasing · 1.7transformer · 0.9tokenization · 0.9FIFO-Diffusion · 0.9training-free control · 0.8principal component analysis · 0.8motion feature extraction · 0.8large language model · 0.8chain-of-contacts · 0.8
YearPublicationVenuePosition
2025 Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space
abstract
Latent Diffusion Models (LDMs) are known to have an unstable generation process, where even small perturbations or shifts in the input noise can lead to significantly different outputs. This hinders their applicability in applications requiring consistent results. In this work, we redesign LDMs to enhance consistency by making them shift-equivariant. While introducing anti-aliasing operations can partially improve shift-equivariance, significant aliasing and inconsistency persist due to the unique challenges in LDMs, including 1) aliasing amplification during VAE training and multiple U-Net inferences, and 2) selfattention modules that inherently lack shift-equivariance. To address these issues, we redesign the attention modules to be shift-equivariant and propose an equivariance loss that effectively suppresses the frequency bandwidth of the features in the continuous domain. The resulting alias-free LDM (AF-LDM) achieves strong shift-equivariance and is also robust to irregular warping. Extensive experiments demonstrate that AF-LDM produces significantly more consistent results than vanilla LDM across various applications, including video editing and image-to-image translation. Code is available at: https://github.com/SingleZombie/AFLDM
Yifan Zhou 0001, Zeqi Xiao, Shuai Yang 0001, Xingang Pan
CVPR2
2025 TokensGen: Harnessing Condensed Tokens for Long Video Generation
abstract
Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In this paper, we propose TokensGen, a novel two-stage framework that leverages condensed tokens to address these issues. Our method decomposes long video generation into three core tasks: (1) inner-clip semantic control, (2) long-term consistency control, and (3) inter-clip smooth transition. First, we train To2V (Token-to-Video), a short video diffusion model guided by text and video tokens, with a Video Tokenizer that condenses short clips into semantically rich tokens. Second, we introduce T2To (Text-to-Token), a video token diffusion transformer that generates all tokens at once, ensuring global consistency across clips. Finally, during inference, an adaptive FIFO-Diffusion strategy seamlessly connects adjacent clips, reducing boundary artifacts and enhancing smooth transitions. Experimental results demonstrate that our approach significantly enhances long-term temporal and content coherence without incurring prohibitive computational overhead. By leveraging condensed tokens and pre-trained short video models, our method provides a scalable, modular solution for long video generation, opening new possibilities for storytelling, cinematic production, and immersive simulations. Please see our project page at https://vicky0522.github.io/tokensgen-webpage/ .
Wenqi Ouyang, Zeqi Xiao, Danni Yang, Yifan Zhou 0001, Shuai Yang 0001, Lei Yang 0045, Jianlou Si, Xingang Pan
ICCV2
2025 Trajectory attention for fine-grained video motion control
abstract
Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixel trajectories for fine-grained camera motion control. Unlike existing methods that often yield imprecise outputs or neglect temporal correlations, our approach possesses a stronger inductive bias that seamlessly injects trajectory information into the video generation process. Importantly, our approach models trajectory attention as an auxiliary branch alongside traditional temporal attention. This design enables the original temporal attention and the trajectory attention to work in synergy, ensuring both precise motion control and new content generation capability, which is critical when the trajectory is only partially available. Experiments on camera motion control for images and videos demonstrate significant improvements in precision and long-range consistency while maintaining high-quality generation. Furthermore, we show that our approach can be extended to other video motion control tasks, such as first-frame-guided video editing, where it excels in maintaining content consistency over large spatial and temporal ranges.
Zeqi Xiao, Wenqi Ouyang, Yifan Zhou 0001, Shuai Yang 0001, Lei Yang 0045, Jianlou Si, Xingang Pan
ICLR1
2025 WorldMem: Long-term Consistent World Simulation with Memory
abstract
World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term consistency, particularly in preserving 3D spatial consistency. In this work, we present WorldMem, a framework that enhances scene generation with a memory bank consisting of memory units that store memory frames and states (e.g., poses and timestamps). By employing state-aware memory attention that effectively extracts relevant information from these memory frames based on their states, our method is capable of accurately reconstructing previously observed scenes, even under significant viewpoint or temporal gaps. Furthermore, by incorporating timestamps into the states, our framework not only models a static world but also captures its dynamic evolution over time, enabling both perception and interaction within the simulated world. Extensive experiments in both virtual and real scenarios validate the effectiveness of our approach.
Zeqi Xiao, Yushi Lan, Yifan Zhou 0001, Wenqi Ouyang, Shuai Yang 0001, Yanhong Zeng, Xingang Pan
NeurIPS1
2025 Position-Guided Point Cloud Panoptic Segmentation Transformer
Zeqi Xiao, Chen Change Loy, Dahua Lin, Jiangmiao Pang
Int. J. Comput. Vis.1
2024 Unified Human-Scene Interaction via Prompted Chain-of-Contacts
abstract
Human-Scene Interaction (HSI) is a vital component of fields like embodied AI and virtual reality. Despite advancements in motion quality and physical plausibility, two pivotal factors, versatile interaction control and the development of a user-friendly interface, require further exploration before the practical application of HSI. This paper presents a unified HSI framework, UniHSI, which supports unified control of diverse interactions through language commands. The framework defines interaction as ``Chain of Contacts (CoC)", representing steps involving human joint-object part pairs. This concept is inspired by the strong correlation between interaction types and corresponding contact regions. Based on the definition, UniHSI constitutes a Large Language Model (LLM) Planner to translate language prompts into task plans in the form of CoC, and a Unified Controller that turns CoC into uniform task execution. To facilitate training and evaluation, we collect a new dataset named ScenePlan that encompasses thousands of task plans generated by LLMs based on diverse scenarios. Comprehensive experiments demonstrate the effectiveness of our framework in versatile task execution and generalizability to real scanned scenes.
Zeqi Xiao, Jingbo Wang 0003, Jinkun Cao, Bo Dai 0002, Dahua Lin, Jiangmiao Pang
ICLR1
2024 CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics
abstract
Enabling humanoid robots to clean rooms has long been a pursued dream within humanoid research communities. However, many tasks require multi-humanoid collaboration, such as carrying large and heavy furniture together. Given the scarcity of motion capture data on multi-humanoid collaboration and the efficiency challenges associated with multi-agent learning, these tasks cannot be straightforwardly addressed using training paradigms designed for single-agent scenarios. In this paper, we introduce **Coo**perative **H**uman-**O**bject **I**nteraction (**CooHOI**), a framework designed to tackle the challenge of multi-humanoid object transportation problem through a two-phase learning paradigm: individual skill learning and subsequent policy transfer. First, a single humanoid character learns to interact with objects through imitation learning from human motion priors. Then, the humanoid learns to collaborate with others by considering the shared dynamics of the manipulated object using centralized training and decentralized execution (CTDE) multi-agent RL algorithms. When one agent interacts with the object, resulting in specific object dynamics changes, the other agents learn to respond appropriately, thereby achieving implicit communication and coordination between teammates. Unlike previous approaches that relied on tracking-based methods for multi-humanoid HOI, CooHOI is inherently efficient, does not depend on motion capture data of multi-humanoid interactions, and can be seamlessly extended to include more participants and a wide range of object types.
Jiawei Gao 0004, Ziqin Wang, Zeqi Xiao, Jingbo Wang 0003, Jinkun Cao, Xiaolin Hu 0001, Si Liu 0001, Jifeng Dai, Jiangmiao Pang
NeurIPS3
2024 Video Diffusion Models are Training-free Motion Interpreter and Controller
abstract
Video generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with training-based paradigms, which, however, demands substantial training resources and necessitates retraining for diverse models. Crucially, these approaches do not explore how video diffusion models encode cross-frame motion information in their features, lacking interpretability and transparency in their effectiveness. To answer this question, this paper introduces a novel perspective to understand, localize, and manipulate motion-aware features in video diffusion models. Through analysis using Principal Component Analysis (PCA), our work discloses that robust motion-aware feature already exists in video diffusion models. We present a new MOtion FeaTure (MOFT) by eliminating content correlation information and filtering motion channels. MOFT provides a distinct set of benefits, including the ability to encode comprehensive motion information with clear interpretability, extraction without the need for training, and generalizability across diverse architectures. Leveraging MOFT, we propose a novel training-free video motion control framework. Our method demonstrates competitive performance in generating natural and faithful motion, providing architecture-agnostic insights and applicability in a variety of downstream tasks.
Zeqi Xiao, Yifan Zhou 0001, Shuai Yang 0001, Xingang Pan
NeurIPS1