VLDB 2026 Research / reviewers in the wild / expert
Xukun Zhou
dblp:344/2377
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0005-5478-1791ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture SynthesisabstractAudio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce coarse gestures, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a novel retrieval-enhanced diffusion framework with three key designs: (1) a Motion Base Construction, which builds a gesture library from the training dataset; (2) a Motion Retrieval Module, employing contrastive learning and momentum distillation for retrieving fine-grained reference poses; and (3) a Precise Control Module, integrating partial masking and stochastic masking to enable flexible and fine-grained control. Experimental evaluations on BEAT2 demonstrate that ExGes reduces Fréchet Gesture Distance by 4.55%and improves motion diversity by 5.3% over EMAGE, with user studies revealing a 71.3% preference for its naturalness and semantic relevance. Xukun Zhou, Fengxin Li, Yan Zhou 0003, Pengfei Wan 0001, Yeying Jin, Hongyuan Zhang 0001, Hongyan Liu 0002, Zhaoxin Fan, Jun He 0008, Xuelong Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Stance Detection for Social Text: Inference-Enhanced Multi-Task Learning with Machine-Annotated SupervisionabstractStance detection, a natural language processing technique, captures user attitudes on controversial social media topics. However, the semantic ambiguity of social texts makes accurate stance determination challenging. Existing annotated data is often domain-specific, resulting in poor model generalization for unseen targets and cross-domain scenarios. To tackle this challenge, we introduce inference tasks related to stance detection as auxiliary tasks and use a multitask learning approach to improve the model’s understanding of textual semantics. Meanwhile, to alleviate the scarcity of annotated data, we use argumentation corpus with abundant resources as the source domain to train the basic stance detection model. We integrate this model with a large-scale framework to facilitate weakly supervised learning for machine annotation of unlabeled social texts, thereby boosting the model’s adaptability across different targets and domains. Experimental results on four Twitter datasets demonstrate that our method can significantly improve the model’s stance discriminative ability and generalization performance. Qiankun Pi, Jicang Lu, Yepeng Sun, Qinlong Fan, Xukun Zhou, Shouxin Shang |
ICASSP | 5 |
| 2025 | Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face AnimationabstractAudio-driven 3D face animation is crucial for live streaming and augmented reality, yet most existing methods focus on specific individuals with predefined speaking styles, limiting adaptability to varied styles. To address this, we introduce MetaFace, a novel methodology for speaking style adaptation based on meta-learning. MetaFace comprises three key components: the Robust Meta Initialization Stage (RMIS) for foundational style adaptation, the Dynamic Relation Mining Neural Process (DRMN) to connect observed and unobserved speaking styles, and a Low-rank Matrix Memory Reduction Approach to optimize model efficiency and style detail learning. These innovations enable MetaFace to significantly outperform existing baselines and set a new state-of-the-art, as demonstrated by our experimental results. Xukun Zhou, Fengxin Li, Ziqiao Peng, Hongyan Liu 0002, Zhaoxin Fan, Jun He 0008 |
ICME | 1 |
| 2024 | BeatDance: A Beat-Based Model-Agnostic Contrastive Learning Framework for Music-Dance RetrievalabstractDance and music are closely related forms of expression, with mutual retrieval between dance videos and music being a fundamental task in various fields like education, art, and sports. However, existing methods often suffer from unnatural generation effects or fail to fully explore the correlation between music and dance. To overcome these challenges, we propose BeatDance, a novel beat-based model-agnostic contrastive learning framework. BeatDance incorporates a Beat-Aware Music-Dance InfoExtractor, a Trans-Temporal Beat Blender, and a Beat-Enhanced Hubness Reducer to improve Music-Dance retrieval performance by utilizing the alignment between music beats and dance movements. We also introduce the Music-Dance (M-D) dataset, a large-scale collection of over 10,000 Music-Dance video pairs for training and testing. Experimental results on the M-D dataset demonstrate the superiority of our method over existing baselines, achieving state-of-the-art performance. The code and dataset are available at https://github.com/XulongT/BeatDance. Kaixing Yang, Xukun Zhou, Xulong Tang, Ran Diao, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan |
ICMR | 2 |
| 2024 | STDG: Semi-Teacher-Student Training Paradigm for Depth-guided One-stage Scene Graph GenerationabstractScene Graph Generation is a critical enabler of environmental comprehension for autonomous robotic systems.Most of existing methods, however, are often thwarted by the intricate dynamics of background complexity, which limits their ability to fully decode the inherent topological information of the environment.Additionally, the wealth of contextual information encapsulated within depth cues is often left untapped, rendering existing approaches less effective.To address these shortcomings, we present STDG, an avant-garde Depth-Guided One-Stage Scene Graph Generation methodology.The innovative architecture of STDG is a triad of custom-built modules: The Depth Guided HHA Representation Generation Module, the Depth Guided Semi-Teaching Network Learning Module, and the Depth Guided Scene Graph Generation Module.This trifecta of modules synergistically harnesses depth information, covering all aspects from depth signal generation and depth feature utilization, to the final scene graph prediction.Importantly, this is achieved without imposing additional computational burden during the inference phase.Experimental results confirm that our method significantly enhances the performance of onestage scene graph generation baselines. Xukun Zhou, Zhenbo Song, Jun He 0008, Hongyan Liu 0002, Zhaoxin Fan |
ICMR | 1 |
| 2024 | Backdoor Attacks with Input-Unique Triggers in NLP
Xukun Zhou, Jiwei Li 0001, Tianwei Zhang 0004, Lingjuan Lyu, Muqiao Yang, Jun He 0008 |
ECML/PKDD (1) | 1 |