EDBT 2026 Demo / reviewers in the wild / expert
Bizhu Wu
dblp:256/7452
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-6783-6561ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FineXtrol: Controllable Motion Generation via Fine-Grained TextabstractRecent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate global 3D coordinate sequences as additional control signals. However, the former often introduces misaligned details and lacks explicit temporal cues, and the latter incurs significant computational cost when converting coordinates to standard motion representations. To address these issues, we propose FineXtrol, a novel control framework for efficient motion generation guided by temporally-aware, precise, user-friendly and fine-grained textual control signals that describe specific body part movements over time. In support of this framework, we design a hierarchical contrastive learning module that encourages the text encoder to produce more discriminative embeddings for our novel control signals, thereby improving motion controllability. Quantitative results show that FineXtrol achieves strong performance in controllable motion generation, while qualitative analysis demonstrates its flexibility in directing specific body part movements. Keming Shen, Bizhu Wu, Junliang Chen 0002, LinLin Shen |
AAAI | 2 |
| 2026 | Ranking-Based Self-Supervised Representation Learning for Skeleton-Based Action RecognitionabstractRecently, researchers have achieved significant results in the skeleton-based action recognition. To better model the skeleton sequences, we drive the encoder to learn more discriminative representations in the self-supervised setting. We find that instead of clustering feature vectors to assign pseudo labels for samples as in DeepCluster, ranking them is a more reasonable, reliable, and efficient way to learn more effective feature representations. With this intuition, we propose a novel self-supervised learning framework,DeepRank. Specifically, we rank triplets of skeleton sequences with the ranking labels, obtained from the relative distances among them. Besides, to deeply mine complementary discriminative information that exists in different modalities of skeleton sequences, we further proposeMulti-ViewDeepRank(MV-DeepRank) to enable encoders to comprehensively learn complementary features from multiple modalities. Extensive experimental results on the NTU RGB+D, NTU RGB+D 120, PKU-MMD I, and PKU-MMD II datasets under various evaluation settings demonstrate the generality, transferability, and superiority of our proposed self-supervised learning frameworks. Notably, our frameworks surpass the previous methods that employ the same backbone networks as ours by at least 1.8% (ST-GCN) and 2.1% (STTFormer) under the finetuning setting. Additionally, DeepRank gains a significant advantage on computational complexities,$O(1)$, over the contrastive learning-based methods,$O(\rm{batch size})$, and the clustering-based methods,$O(\rm{number of clusters})$. Bizhu Wu, Junliang Chen 0002, Jinheng Xie, Qiufu Li, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
IEEE Trans. Multim. | 1 |
| 2025 | MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple GranularitiesabstractRecent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
CVPR | 1 |
| 2025 | FineMotion: A Dataset and Benchmark with Both Spatial and Temporal Annotation for Fine-Grained Motion Generation and Editing
Bizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
ICCV | 1 |
| 2024 | OMG: Occlusion-Friendly Personalized Multi-concept Generation in Diffusion Models
Zhe Kong, Yong Zhang 0034, Tianyu Yang 0003, Tao Wang 0052, Kaihao Zhang, Bizhu Wu, Guanying Chen, Wei Liu 0005, Wenhan Luo |
ECCV (31) | 6 |
| 2023 | Weakly Supervised Pedestrian Segmentation for Person Re-IdentificationabstractPerson re-identification (RelD) is an important problem in intelligent surveillance and public security. Among all the solutions to this problem, existing mask-based methods first use a well-pretrained segmentation model to generate a foreground mask, in order to exclude the background from ReID. Then they perform the RelD task directly on the segmented pedestrian image. However, such a process requires extra datasets with pixel-level semantic labels. In this paper, we propose a Weakly Supervised Pedestrian Segmentation (WSPS) framework to produce the foreground mask directly from the RelD datasets. In contrast, our WSPS only requires image-level subject ID labels. To better utilize the pedestrian mask, we also propose the Image Synthesis Augmentation (ISA) technique to further augment the dataset. Experiments show that the features learned from our proposed framework are robust and discriminative. Compared with the baseline, the mAP of our framework is about 4.4%, 11.7%, and 4.0% higher on three widely used datasets including Market-1501, CUHK03, and MSMT17. The code will be available soon. Ziqi Jin, Jinheng Xie, Bizhu Wu, LinLin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Frequency-driven Imperceptible Adversarial Attack on Semantic SimilarityabstractCurrent adversarial attack research reveals the vulnerability of learning-based classifiers against carefully crafted perturbations. However, most existing attack methods have inherent limitations in cross-dataset generalization as they rely on a classification layer with a closed set of categories. Furthermore, the perturbations generated by these methods may appear in regions easily perceptible to the human visual system (HVS). To circumvent the former problem, we propose a novel algorithm that attacks semantic similarity on feature representations. In this way, we are able to fool classifiers without limiting attacks to a specific dataset. For imperceptibility, we introduce the low-frequency constraint to limit perturbations within high-frequency components, ensuring perceptual similarity between adversarial examples and originals. Extensive experiments on three datasets (CIFAR-10, CIFAR-100, and ImageNet-1K) and three public online platforms indicate that our attack can yield misleading and transferable adversarial examples across architectures and datasets. Additionally, visualization results and quantitative performance (in terms of four different metrics) show that the proposed algorithm generates more imperceptible perturbations than the state-of-the-art methods. Code is made available at https://github.com/LinQinLiang/SSAH-adversarial-attack. Qinliang Lin, Weicheng Xie 0001, Bizhu Wu, Jinheng Xie, LinLin Shen |
CVPR | 4 |
| 2022 | Which One is Better? Self-supervised Temporal Coherence Learning for Skeleton Based Action RecognitionabstractRecently, researchers have achieved significant results in the skeleton based action recognition task. To better model the skeleton sequences, existing methods learned the feature representations in the self-supervised setting by solving pretext tasks, such as predicting the order of a shuffled skeleton sequence or verifying whether a given skeleton sequence is shuffled or not. However, these pretext tasks are either too challenging or too easy for the encoder to obtain a proper skeleton representation for action recognition. Therefore, we propose a novel self-pretraining pretext task, Which One Is Better (WOIB), to identify which one is more temporally coherent, given two shuffled skeleton sequences. Experiments on the NTU RGB+D, NTU RGB+D 120, and Kinetics-Skeleton datasets with different network architectures show significant improvements in recognition accuracy, demonstrating that such a well-designed pretext task is general and able to drive the encoder to learn more discriminative representations. Bizhu Wu, Mingyan Wu, Haoqin Ji, LinLin Shen |
IJCB | 1 |