VLDB 2026 Research / reviewers in the wild / expert
Junbo Wang 0003
dblp:17/5974-3
· DBLP profile ↗
13ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0006-9955-7838ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided AlignmentabstractText-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and realistic motion. However, there exists a misalignment between text and motion distributions in diffusion models, which leads to semantically inconsistent or low-quality motions. To address this limitation, we propose Reward-guided sampling Alignment (ReAlign), comprising a step-aware reward model to assess alignment quality during the denoising sampling and a reward-guided strategy that directs the diffusion process toward an optimally aligned distribution. This reward model integrates step-aware tokens and combines a text-aligned module for semantic consistency and a motion-aligned module for realism, refining noisy motions at each timestep to balance probability density and alignment. Extensive experiments of both motion generation and retrieval tasks demonstrate that our approach significantly improves text-motion alignment and motion quality compared to existing state-of-the-art methods. Wanjiang Weng, Xiaofeng Tan 0001, Junbo Wang 0003, Guosen Xie, Pan Zhou 0002, Hongsong Wang 0001 |
AAAI | 3 |
| 2026 | Foundation Model for Skeleton-Based Human Action UnderstandingabstractHuman action understanding serves as a foundational pillar in the field of intelligent motion perception.Skeletons serve as a modality- and device-agnostic representation for human modeling, and skeleton-based action understanding has potential applications in humanoid robot control and interaction. However, existing works often lack the scalability and generalization required to handle diverse action understanding tasks. There is no skeleton foundation model that can be adapted to a wide range of action understanding tasks. This paper presents a Unified Skeleton-based Dense Representation Learning (USDRL) framework, which serves as a foundational model for skeleton-based human action understanding. USDRL consists of a Transformer-based Dense Spatio-Temporal Encoder (DSTE), Multi-Grained Feature Decorrelation (MG-FD), and Multi-Perspective Consistency Training (MPCT). The DSTE module adopts two parallel streams to learn temporal dynamic and spatial structure features. The MG-FD module collaboratively performs feature decorrelation across temporal, spatial, and instance domains to reduce dimensional redundancy and enhance information extraction. The MPCT module employs both multi-view and multi-modal self-supervised consistency training. The former enhances the learning of high-level semantics and mitigates the impact of low-level discrepancies, while the latter effectively facilitates the learning of informative multimodal features. We perform extensive experiments on 25 benchmarks across across 9 skeleton-based action understanding tasks, covering coarse prediction, dense prediction, and transferred prediction. Our approach significantly outperforms the current state-of-the-art methods. We hope that this work would broaden the scope of research in skeleton-based action understanding and encourage more attention to dense prediction tasks. Hongsong Wang 0001, Wanjiang Weng, Junbo Wang 0003, Fang Zhao 0006, Guosen Xie, Xin Geng 0001, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationabstractContrastive learning has achieved great success in skeleton-based representation learning recently. However, the prevailing methods are predominantly negative-based, necessitating additional momentum encoder and memory bank to get negative samples, which increases the difficulty of model training. Furthermore, these methods primarily concentrate on learning a global representation for recognition and retrieval tasks, while overlooking the rich and detailed local representations that are crucial for dense prediction tasks. To alleviate these issues, we introduce a Unified Skeleton-based Dense Representation Learning framework based on feature decorrelation, called USDRL, which employs feature decorrelation across temporal, spatial, and instance domains in a multi-grained manner to reduce redundancy among dimensions of the representations to maximize information extraction from features. Additionally, we design a Dense Spatio-Temporal Encoder (DSTE) to capture fine-grained action representations effectively, thereby enhancing the performance of dense prediction tasks. Comprehensive experiments, conducted on the benchmarks NTU-60, NTU-120, PKU-MMD I, and PKU-MMD II, across diverse downstream tasks including action recognition, action retrieval, and action detection, conclusively demonstrate that our approach significantly outperforms the current state-of-the-art (SOTA) approaches. Wanjiang Weng, Hongsong Wang 0001, Junbo Wang 0003, Guosen Xie |
AAAI | 3 |
| 2025 | PTSR: A Unified Patch Tokenization, Selection and Representation Framework for Efficient Micro-expression RecognitionabstractMicro-expression recognition is a challenging task of identifying hidden emotion, as micro-expressions have brief durations and involve small-scale facial muscle movements. Although deep learning-based methods, especially transformer-based methods, have achieved impressive performance in this task, these methods exhibit high computational complexity and struggle to learn effective representations in the context of typically small-scale micro-expression datasets, due to the excess of tokens in the multi-head self-attention. Moreover, most existing methods do not differentiate the importance of local features, especially in micro-expression recognition with subtle changes. Therefore, we propose a novel unified Patch Tokenization, Selection and Representation framework (PTSR) with vision Transformer for micro-expression recognition. Specifically, PTSR first presents a dual norm shifted patch tokenization (DNSPT) module to learn spatial relations between neighboring pixels of the face region, which is implemented by elaborating spatial transformation and dual norm projection. Then, we employ a local-global attention module (LAM) to extract the local-global image feature, incorporating a dynamic token selection module (DTSM) to select important patches/tokens, thereby capturing more discriminative representations for the input clip. Extensive experiments are conducted on 4 widely used public datasets, i.e., CASME II, SAMM, SMIC, CAS(ME)3, and the experimental results indicate that our method can achieve clear performance improvements over the state-of-the-art methods, such as 8.37% improvement on the CAS(ME)3 dataset in terms of UF1 and 3.1% improvement on the SMIC dataset in terms of UAR metric. Liangyu Fu, Junbo Wang 0003, Qiangguo Jin, Yining Zhu, Hongsong Wang 0001, Kun Hu 0008 |
ICMR | 2 |
| 2025 | MirrorDiff: Learning Mirror Diffusion for Image Captioning via RegenerationabstractRecently, diffusion models which have achieved promising progress in text-to-image generation generally have also been generally explored for image captioning. However, these diffusion-based image captioning methods usually suffer from semantic inconsistency between image content and textual description, thus producing lagging results compared with Auto-Regressive (AR) ones. To this end, in this paper, we propose a novel dual diffusion-based framework namely MirrorDiff, to achieve semantic consistency with a symmetric image-to-text-to-image generation model, which acts like a mirror that maps the original input image into a regenerated image via the generated caption. Specifically, it first utilizes both pre-trained image encoder and text encoder to obtain image representation and textual representation respectively, then forwards the image representation and the noisy textual representation into a continuous diffusion model to output an intermediate sentence. To semantically align the intermediate sentence with the input image, a diffusion-based visual regenerator is employed to regenerate the input image conditioned on the intermediate sentence, resulting in a proposed visual regeneration loss. Different from most existing image captioning methods, MirrorDiff is a plug-and-play framework which can be plugged into many previous image captioning methods, and further evaluate the generated sentence via the visual similarity between the input image and the regenerated image. Extensive experiments on the MS COCO dataset show that our method achieves obvious improvements over state-of-the-art diffusion-based methods, up to 127.9 on CIDEr, and achieves competitive performance on multiple evaluation metrics over the auto-regressive methods trained on larger-scale datasets. Junbo Wang 0003, Liangyu Fu, Yining Zhu, Qiangguo Jin, Hongsong Wang 0001, Kun Hu 0008 |
ICMR | 1 |
| 2025 | DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image Captioning
Liangyu Fu, Junbo Wang 0003, Qiangguo Jin, Hongsong Wang 0001, Jing Ya, Linjiang Huang, Jiangbin Zheng 0001, Zhiyong Wang 0001 |
ACM Multimedia | 2 |
| 2024 | SPK: Semantic and Positional Knowledge for Zero-Shot Referring Expression Comprehension
Zetao Du, Junbo Wang 0003, Yan Huang 0008, Liang Wang 0001 |
ICPR (30) | 3 |
| 2020 | Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchabstractText-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting visual contents corresponding to the human description is the key to this cross-modal matching problem. Moreover, correlated images and descriptions involve different granularities of semantic relevance, which is usually ignored in previous methods. To exploit the multilevel corresponding visual contents, we propose a pose-guided multi-granularity attention network (PMA). Firstly, we propose a coarse alignment network (CA) to select the related image regions to the global description by a similarity-based attention. To further capture the phrase-related visual body part, a fine-grained alignment network (FA) is proposed, which employs pose information to learn latent semantic alignment between visual body part and textual noun phrase. To verify the effectiveness of our model, we perform extensive experiments on the CUHK Person Description Dataset (CUHK-PEDES) which is currently the only available dataset for text-based person search. Experimental results show that our approach outperforms the state-of-the-art methods by 15 % in terms of the top-1 metric. Ya Jing, Chenyang Si, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
AAAI | 3 |
| 2020 | Relational graph neural network for situation recognition
Ya Jing, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 2 |
| 2020 | Learning visual relationship and context-aware attention for image captioning
Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Zhiyong Wang 0001, David Dagan Feng, Tieniu Tan |
Pattern Recognit. | 1 |
| 2019 | Stacked Memory Network for Video SummarizationabstractIn recent years, supervised video summarization has achieved promising progress with various recurrent neural networks (RNNs) based methods, which treats video summarization as a sequence-to-sequence learning problem to exploit temporal dependency among video frames across variable ranges. However, RNN has limitations in modelling the long-term temporal dependency for summarizing videos with thousands of frames due to the restricted memory storage unit. Therefore, in this paper we propose a stacked memory network called SMN to explicitly model the long dependency among video frames so that redundancy could be minimized in the video summaries produced. Our proposed SMN consists of two key components: Long Short-Term Memory (LSTM) layer and memory layer, where each LSTM layer is augmented with an external memory layer. In particular, we stack multiple LSTM layers and memory layers hierarchically to integrate the learned representation from prior layers. By combining the hidden states of the LSTM layers and the read representations of the memory layers, our SMN is able to derive more accurate video summaries for individual video frames. Compared with the existing RNN based methods, our SMN is particularly good at capturing long temporal dependency among frames with few additional training parameters. Experimental results on two widely used public benchmark datasets: SumMe and TVsum, demonstrate that our proposed model is able to clearly outperform a number of state-of-the-art ones under various settings. Junbo Wang 0003, Wei Wang 0115, Zhiyong Wang 0001, Liang Wang 0001, David Dagan Feng, Tieniu Tan |
ACM Multimedia | 1 |
| 2018 | M3: Multimodal Memory Modelling for Video CaptioningabstractVideo captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, video captioning has made great progress. However, learning an effective mapping from the visual sequence space to the language space is still a challenging problem due to the long-term multimodal dependency modelling and semantic misalignment. Inspired by the facts that memory modelling poses potential advantages to long-term sequential problems [35] and working memory is the key factor of visual attention [33], we propose a Multimodal Memory Model (M3) to describe videos, which builds a visual and textual shared memory to model the long-term visual-textual dependency and further guide visual attention on described visual targets to solve visual-textual alignments. Specifically, similar to [10], the proposed M3 attaches an external memory to store and retrieve both visual and textual contents by interacting with video and sentence with multiple read and write operations. To evaluate the proposed model, we perform experiments on two public datasets: MSVD and MSR-VTT. The experimental results demonstrate that our method outperforms most of the state-of-the-art methods in terms of BLEU and METEOR. Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
CVPR | 1 |
| 2018 | Hierarchical Memory Modelling for Video CaptioningabstractTranslating videos into natural language sentences has drawn much attention recently. The framework of combining visual attention with Long Short-Term Memory (LSTM) based text decoder has achieved much progress. However, the vision-language translation still remains unsolved due to the semantic gap and misalignment between video content and described semantic concept. In this paper, we propose a Hierarchical Memory Model (HMM) - a novel deep video captioning architecture which unifies a textual memory, a visual memory and an attribute memory in a hierarchical way. These memories can guide attention for efficient video representation extraction and semantic attribute selection in addition to modelling the long-term dependency for video sequence and sentences, respectively. Compared with traditional vision-based text decoder, the proposed attribute-based text decoder can largely reduce the semantic discrepancy between video and sentence. To prove the effectiveness of the proposed model, we perform extensive experiments on two public benchmark datasets: MSVD and MSR-VTT. Experiments show that our model not only can discover appropriate video representation and semantic attributes but also can achieve comparable or superior performances than state-of-the-art methods on these datasets. Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
ACM Multimedia | 1 |