Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xunzhi Xiang

dblp:374/8077 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0006-9629-0410ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Image recognition and object detection · 36% Transfer learning and domain adaptation · 18% Generative modeling · 18%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection › few-shot object detection
cross-domain few-shot object detection
0.912025
DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object Detection · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
domain generalization
0.912025
DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object Detection · NeurIPS 2025
Computer vision › Image recognition and object detection › object detection
few-shot object detection
0.912025
DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object Detection · NeurIPS 2025
Visual content generation and editing › image animation
human image animation
0.912025
ReMask-Animate: Refined Character Image Animation Using Mask-Guided Adapters · AAAI 2025
Visual content generation and editing › video editing
video customization
0.912025
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization · SIGGRAPH Asia 2025
Visual content generation and editing
video generation
0.912025
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization · SIGGRAPH Asia 2025
Computer vision › Face, body and person analysis › face recognition
identity preservation
0.312025
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization · SIGGRAPH Asia 2025
Natural language and speech › Language models and text generation
large language model
0.312025
Human Motion Video Generation: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.6multimodal identity fusion · 1.7adaptive motion learning · 1.7vision foundation model · 0.9routing · 0.9mixture of experts · 0.9mask-guided adapters · 0.9large language model · 0.9cross-condition regional fusion · 0.9
YearPublicationVenuePosition
2025 ReMask-Animate: Refined Character Image Animation Using Mask-Guided Adapters
abstract
Pose-controlled human video generation is of significant interest and finds extensive applications in areas such as automated advertising and content creation on social media platforms. While existing methods employing pose sequences and reference images for human image animation have exhibited notable performance, they tend to encounter issues such as specific region blurring, background sharpening, and decreased identity consistency. In this paper, we introduce ReMask-Animate, which utilizes masks as additional priors to guide the model's local visual attention to specific areas, thereby alleviating feature confusion between different regions of the image. Three distinct mask-guided adapters are designed for cross-condition regional fusion of hand and face pose features, mitigating feature confusion between the foreground and background, and enhancing the visual consistency of character identity. Moreover, these lightweight adapters introduce minimal computational overhead and can be seamlessly integrated into specific layers of the backbone architecture. Extensive experiments show that our method outperforms state-of-the-art methods on five metrics in public datasets. Additionally, qualitative evaluations highlight a significant improvement in the quality of generated videos, demonstrating our approach's superiority.
Xunzhi Xiang, Haiwei Xue, Zonghong Dai, Minglei Li 0001, Ye Yue, Fei Ma 0006, Weijiang Yu, Heng Chang, F. Richard Yu
AAAI1
2025 DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object Detection
abstract
Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to generalize to unseen domains by leveraging a few annotated samples of the target domain, requiring models to exhibit both strong generalization and localization capabilities. However, existing well-trained detectors typically have strong localization capabilities but lack generalization, whereas vision foundation models (VFMs) generally exhibit better generalization but lack accurate localization capabilities. In this paper, we propose a novel Mixture-of-Experts (MoE) structure that integrates the detector's localization capability and the VFM's generalization by using VFM features to improve detector features. Specifically, we propose Expert-wise Router (ER) that selects the most relevant VFM experts for each backbone layer, and Region-wise Router (RR) that emphasizes foreground and suppress background. To bridge representation gaps, we further propose Shared Expert Projection (SEP) module and Private Expert Projection (PEP) module, which align VFM features to the detector feature space while decoupling shared image feature from private image feature in the VFM feature map. Finally, we propose MoE module to transfer the VFM’s generalization to the detector without altering the detector original architecture. Furthermore, our method extend well-trained detectors for detecting novel classes in unseen domains without re-training on the base classes. Experimental results on multiple cross-domain datasets validate the effectiveness of our method.
Changhan Liu, Xunzhi Xiang, Zixuan Duan, Yang Gao 0001
NeurIPS2
2025 Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
abstract
Video identity customization seeks to synthesize realistic, temporally coherent videos of a specific subject, given a single reference image and a text prompt. This task presents two core challenges: (1) maintaining identity consistency while aligning with the described appearance and actions, and (2) generating natural, fluid motion without unrealistic stiffness. To address these challenges, we introduce Proteus-ID, a novel diffusion-based framework for identity-consistent and motion-coherent video customization. First, we propose a Multimodal Identity Fusion (MIF) module that unifies visual and textual cues into a joint identity representation using a Q-Former, providing coherent guidance to the diffusion model and eliminating modality imbalance. Second, we present a Time-Aware Identity Injection (TAII) mechanism that dynamically modulates identity conditioning across denoising steps, improving fine-detail reconstruction. Third, we propose Adaptive Motion Learning (AML), a motion-aware optimization strategy that reweights training loss based on optical-flow-derived motion heatmaps, enhancing motion realism without requiring additional inputs. To support this task, we construct Proteus-Bench, a high-quality dataset comprising 200K curated clips for training and 150 individuals from diverse professions and ethnicities for evaluation. Extensive experiments demonstrate that Proteus-ID outperforms prior methods in identity preservation, text alignment, and motion quality, establishing a new benchmark for video identity customization.
Guiyu Zhang, Zijian Jiang, Xunzhi Xiang, Jingjing Qian, Shaoshuai Shi, Li Jiang 0009
SIGGRAPH Asia4
2025 Human Motion Video Generation: A Survey
abstract
Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans.
Haiwei Xue, Xiangyang Luo 0002, Zhanghao Hu, Xin Zhang 0169, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li 0001, Jian Yang 0003, Fei Ma 0006, Zhiyong Wu 0001, Changpeng Yang, Zonghong Dai, F. Richard Yu
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 A Neural Architecture Predictor based on GNN-Enhanced Transformer
abstract
Neural architecture performance predictor is an efficient approach for architecture estimation in Neural Architecture Search (NAS). However, existing predictors based on Graph Neural Networks (GNNs) are deficient in modeling long-range interactions between operation nodes and prone to the problem of over-smoothing, which limits their ability to learn neural architecture representation. Furthermore, some Transformer-based predictors use simple position encodings to improve performance via self-attention mechanism, but they fail to fully exploit the subgraph structure information of the graph. To solve this problem, we propose a novel method to enhance the graph representation of neural architectures by combining GNNs and Transformer blocks. We evaluate the effectiveness of our predictor on NAS-Bench-101 and NAS-bench-201 benchmarks, the discovered architecture on DARTS search space achieves an accuracy of 97.61% on CIFAR-10 dataset, which outperforms traditional position encoding methods such as adjacency and Laplacian matrices. The code of our work is available at \url{https://github.com/GNET}.
Xunzhi Xiang, Kun Jing, Jungang Xu
AISTATS1
2024 Feature Activation-Driven Zero-Shot NAS: A Contrastive Learning Framework
Di Wang 0053, Xunzhi Xiang, Kun Jing, Jungang Xu
ICANN (1)2
2024 Improving Image Captioning with Image Concepts of Words
Xunzhi Xiang, Kun Jing, Jungang Xu, Yingfei Sun
KSEM (2)2