VLDB 2026 Research / reviewers in the wild / expert
Hongsong Wang 0001
dblp:181/4564-1
· DBLP profile ↗
32ranked-venue papers
15as first author
27since 2021 · last 2026
0000-0002-9464-1778ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 11 first-author · 19 since 2021Artificial intelligence and machine learning · 16 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided AlignmentabstractText-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and realistic motion. However, there exists a misalignment between text and motion distributions in diffusion models, which leads to semantically inconsistent or low-quality motions. To address this limitation, we propose Reward-guided sampling Alignment (ReAlign), comprising a step-aware reward model to assess alignment quality during the denoising sampling and a reward-guided strategy that directs the diffusion process toward an optimally aligned distribution. This reward model integrates step-aware tokens and combines a text-aligned module for semantic consistency and a motion-aligned module for realism, refining noisy motions at each timestep to balance probability density and alignment. Extensive experiments of both motion generation and retrieval tasks demonstrate that our approach significantly improves text-motion alignment and motion quality compared to existing state-of-the-art methods. Wanjiang Weng, Xiaofeng Tan 0001, Junbo Wang 0003, Guosen Xie, Pan Zhou 0002, Hongsong Wang 0001 |
AAAI | 6 |
| 2026 | Foundation Model for Skeleton-Based Human Action UnderstandingabstractHuman action understanding serves as a foundational pillar in the field of intelligent motion perception.Skeletons serve as a modality- and device-agnostic representation for human modeling, and skeleton-based action understanding has potential applications in humanoid robot control and interaction. However, existing works often lack the scalability and generalization required to handle diverse action understanding tasks. There is no skeleton foundation model that can be adapted to a wide range of action understanding tasks. This paper presents a Unified Skeleton-based Dense Representation Learning (USDRL) framework, which serves as a foundational model for skeleton-based human action understanding. USDRL consists of a Transformer-based Dense Spatio-Temporal Encoder (DSTE), Multi-Grained Feature Decorrelation (MG-FD), and Multi-Perspective Consistency Training (MPCT). The DSTE module adopts two parallel streams to learn temporal dynamic and spatial structure features. The MG-FD module collaboratively performs feature decorrelation across temporal, spatial, and instance domains to reduce dimensional redundancy and enhance information extraction. The MPCT module employs both multi-view and multi-modal self-supervised consistency training. The former enhances the learning of high-level semantics and mitigates the impact of low-level discrepancies, while the latter effectively facilitates the learning of informative multimodal features. We perform extensive experiments on 25 benchmarks across across 9 skeleton-based action understanding tasks, covering coarse prediction, dense prediction, and transferred prediction. Our approach significantly outperforms the current state-of-the-art methods. We hope that this work would broaden the scope of research in skeleton-based action understanding and encourage more attention to dense prediction tasks. Hongsong Wang 0001, Wanjiang Weng, Junbo Wang 0003, Fang Zhao 0006, Guosen Xie, Xin Geng 0001, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Zero-shot skeleton-based action recognition with dual visual-text alignment
Jidong Kuang, Hongsong Wang 0001, Chaolei Han 0001, Yang Zhang 0002, Jie Gui |
Pattern Recognit. | 2 |
| 2026 | Efficient Neural Architecture Search for brain-inspired spiking neural network
Peilin Lai, Yang Zhang 0002, Weizhao He, Hongsong Wang 0001 |
Pattern Recognit. | 5 |
| 2026 | Efficient Diffusion-Based 3D Human Pose Estimation With Hierarchical Temporal PruningabstractDiffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial computational cost. In this paper, we propose an efficient diffusion-based 3D human pose estimation framework with a Hierarchical Temporal Pruning (HTP) strategy, which dynamically prunes redundant pose tokens across both frame and semantic levels while preserving critical motion dynamics. HTP operates in a staged, top-down manner: (1) Temporal Correlation-Enhanced Pruning (TCEP) identifies essential frames by analyzing inter-frame motion correlations through adaptive temporal graph construction; (2) Sparse-Focused Temporal MHSA (SFT MHSA) leverages the resulting frame-level sparsity to reduce attention computation, focusing on motion-relevant tokens; and (3) Mask-Guided Pose Token Pruner (MGPTP) performs fine-grained semantic pruning via clustering, retaining only the most informative pose tokens. Experiments on Human3.6M and MPI-INF-3DHP show that HTP reduces training MACs by 38.5%, inference MACs by 56.8%, and improves inference speed by an average of 81.1% compared to prior diffusion-based methods, while achieving state-of-the-art performance. Yuquan Bi, Hongsong Wang 0001, Xinli Shi, Zhipeng Gui, Jie Gui, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Data-Free Class-Incremental Gesture Recognition With Prototype-Guided Pseudo-Feature ReplayabstractGesture recognition is an important research area in the field of computer vision. Most existing efforts focus on close-set scenarios, thereby limiting the capacity to effectively handle unseen or novel gestures. We aim to address class-incremental gesture recognition, which entails the ability to accommodate new and previously unseen gestures over time. Specifically, we introduce a Prototype-Guided Pseudo Feature Replay framework for data-free class-incremental learning. This framework comprises four components: Pseudo Feature Generation with Batch Prototypes (PFGBP), Variational Prototype Replay for old classes, Truncated Cross-Entropy for new classes, and Continual Classifier Re-Training. To tackle the issue of catastrophic forgetting, the PFGBP dynamically generates a diversity of pseudo features in an online manner, leveraging class prototypes of old classes along with batch class prototypes of new classes. Furthermore, the Variational Prototype Replay enforces consistency between the classifier's weights and the prototypes of old classes, leveraging class prototypes and covariance matrices to enhance robustness and generalization capabilities. The Truncated Cross-Entropy mitigates the impact of domain differences of the classifier caused by pseudo features. Finally, the Continual Classifier Re-Training training strategy is designed to prevent overfitting to new classes and ensure the stability of features extracted from old classes. Extensive experiments conducted on two widely used gesture recognition datasets, namely SHREC 2017 3D and EgoGesture 3D, demonstrate that our approach outperforms existing state-of-the-art methods by 11.8% and 12.8% in terms of mean global accuracy, respectively. The code is available on https://github.com/sunao-101/PGPFR-3/. Hongsong Wang 0001, Jie Gui, Liang Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Location Matters: Frequency-Spatial Dual-Space Adaptation for Cross-Domain Few-Shot SegmentationabstractCurrent cross-domain few-shot semantic segmentation (CD-FSS) methods tend to overlook a fundamental yet domain-agnostic prior: the spatial correspondence between support and query images driven by the task itself. Unlike semantic similarity, this spatial correlation arises from the consistent structural layout of foreground objects across domains. To exploit this structural prior, we propose a novel frequency-spatial dual space adaptation (FDSA) framework, to learn domain-invariant structures and task-specific priors by jointly suppressing domain-specific redundancy in frequency domain and reinforcing geometric priors in spatial domain. Specifically, FDSA consists of two sequential modules, i.e., the frequency structural adapter (FSA) and the spatial geometry adapter (SGA). FSA performs image modulation in the frequency domain by emphasizing low-frequency foreground semantics and attenuating high-frequency noise, thus maintaining structural integrity of these input images. By contrast, SGA leverages handcrafted local descriptors to extract keypoints from both support and query images, generating Gaussian-based geometric priors that highlight desirable aligned regions. Additionally, we introduce spatial-guided SAM refinement (SSR) to extend our spatial geometric prior into the Segment Anything Model (SAM). SSR generates a soft Gaussian point prompt centered on the coarse mask, enabling SAM to refine segmentation masks without manual intervention. This integration effectively bridges task-specific localization with high-quality segmentation. Extensive experiments on four standard CD-FSS benchmarks demonstrate that our method achieves new state-of-the-art performance. Code is available at https://github.com/CVL-hub/FDSA.git. Guolei Sun, Yong Li 0032, Hongsong Wang 0001, Xiangbo Shu, Guosen Xie |
IEEE Trans. Image Process. | 4 |
| 2025 | SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot SegmentationabstractThe primary challenge of cross-domain few-shot segmentation (CD-FSS) is the domain disparity between the training and inference phases, which can exist in either the input data or the target classes. Previous models struggle to learn feature representations that generalize to various unknown domains from limited training domain samples. In contrast, the large-scale visual model SAM, pre-trained on tens of millions of images from various domains and classes, possesses excellent generalizability. In this work, we propose a SAM-aware graph prompt reasoning network (GPRN) that fully leverages SAM to guide CD-FSS feature representation learning and improve prediction accuracy. Specifically, we propose a SAM-aware prompt initialization module (SPI) to transform the masks generated by SAM into visual prompts enriched with high-level semantic information. Since SAM tends to divide an object into many sub-regions, this may lead to visual prompts representing the same semantic object having inconsistent or fragmented features. We further propose a graph prompt reasoning (GPR) module that constructs a graph among visual prompts to reason about their interrelationships and enable each visual prompt to aggregate information from similar prompts, thus achieving global semantic consistency. Subsequently, each visual prompt embeds its semantic information into the corresponding mask region to assist in feature representation learning. To refine the segmentation mask during testing, we also design a non-parameter adaptive point selection module (APS) to select representative point prompts from query predictions and feed them back to SAM to refine inaccurate segmentation results. Experiments on four standard CD-FSS datasets demonstrate that our method establishes new state-of-the-art results. Shi-Feng Peng, Guolei Sun, Yong Li 0032, Hongsong Wang 0001, Guosen Xie |
AAAI | 4 |
| 2025 | Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is essential for computer vision and multimedia research. Existing VAD methods utilize either reconstruction-based or prediction-based frameworks. The former excels at detecting irregular patterns or structures, whereas the latter is capable of spotting abnormal deviations or trends. We address pose-based video anomaly detection and introduce a novel framework called Dual Conditioned Motion Diffusion (DCMD), which enjoys the advantages of both approaches. The DCMD integrates conditioned motion and conditioned embedding to comprehensively utilize the pose characteristics and latent semantics of observed movements, respectively. In the reverse diffusion process, a motion transformer is proposed to capture potential correlations from multi-layered characteristics within the spectrum space of human motion. To enhance the discriminability between normal and abnormal instances, we design a novel United Association Discrepancy (UAD) regularization that primarily relies on a Gaussian kernel-based time association and a self-attention-based global association. Finally, a mask completion strategy is introduced during the inference stage of the reverse diffusion process to enhance the utilization of conditioned motion for the prediction branch of anomaly detection. Extensive experiments conducted on four datasets demonstrate that our method dramatically outperforms state-of-the-art methods and exhibits superior generalization performance. Hongsong Wang 0001, Andi Xu, Pinle Ding, Jie Gui |
AAAI | 1 |
| 2025 | USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationabstractContrastive learning has achieved great success in skeleton-based representation learning recently. However, the prevailing methods are predominantly negative-based, necessitating additional momentum encoder and memory bank to get negative samples, which increases the difficulty of model training. Furthermore, these methods primarily concentrate on learning a global representation for recognition and retrieval tasks, while overlooking the rich and detailed local representations that are crucial for dense prediction tasks. To alleviate these issues, we introduce a Unified Skeleton-based Dense Representation Learning framework based on feature decorrelation, called USDRL, which employs feature decorrelation across temporal, spatial, and instance domains in a multi-grained manner to reduce redundancy among dimensions of the representations to maximize information extraction from features. Additionally, we design a Dense Spatio-Temporal Encoder (DSTE) to capture fine-grained action representations effectively, thereby enhancing the performance of dense prediction tasks. Comprehensive experiments, conducted on the benchmarks NTU-60, NTU-120, PKU-MMD I, and PKU-MMD II, across diverse downstream tasks including action recognition, action retrieval, and action detection, conclusively demonstrate that our approach significantly outperforms the current state-of-the-art (SOTA) approaches. Wanjiang Weng, Hongsong Wang 0001, Junbo Wang 0003, Guosen Xie |
AAAI | 2 |
| 2025 | Heterogeneous Skeleton-Based Action Representation Learning
Hongsong Wang 0001, Jidong Kuang, Jie Gui |
CVPR | 1 |
| 2025 | LOTA: Bit-Planes Guided AI-Generated Image Detection
Hongsong Wang 0001, Renxi Cheng, Yang Zhang 0002, Chaolei Han 0001, Jie Gui |
ICCV | 1 |
| 2025 | PTSR: A Unified Patch Tokenization, Selection and Representation Framework for Efficient Micro-expression RecognitionabstractMicro-expression recognition is a challenging task of identifying hidden emotion, as micro-expressions have brief durations and involve small-scale facial muscle movements. Although deep learning-based methods, especially transformer-based methods, have achieved impressive performance in this task, these methods exhibit high computational complexity and struggle to learn effective representations in the context of typically small-scale micro-expression datasets, due to the excess of tokens in the multi-head self-attention. Moreover, most existing methods do not differentiate the importance of local features, especially in micro-expression recognition with subtle changes. Therefore, we propose a novel unified Patch Tokenization, Selection and Representation framework (PTSR) with vision Transformer for micro-expression recognition. Specifically, PTSR first presents a dual norm shifted patch tokenization (DNSPT) module to learn spatial relations between neighboring pixels of the face region, which is implemented by elaborating spatial transformation and dual norm projection. Then, we employ a local-global attention module (LAM) to extract the local-global image feature, incorporating a dynamic token selection module (DTSM) to select important patches/tokens, thereby capturing more discriminative representations for the input clip. Extensive experiments are conducted on 4 widely used public datasets, i.e., CASME II, SAMM, SMIC, CAS(ME)3, and the experimental results indicate that our method can achieve clear performance improvements over the state-of-the-art methods, such as 8.37% improvement on the CAS(ME)3 dataset in terms of UF1 and 3.1% improvement on the SMIC dataset in terms of UAR metric. Liangyu Fu, Junbo Wang 0003, Qiangguo Jin, Yining Zhu, Hongsong Wang 0001, Kun Hu 0008 |
ICMR | 5 |
| 2025 | MirrorDiff: Learning Mirror Diffusion for Image Captioning via RegenerationabstractRecently, diffusion models which have achieved promising progress in text-to-image generation generally have also been generally explored for image captioning. However, these diffusion-based image captioning methods usually suffer from semantic inconsistency between image content and textual description, thus producing lagging results compared with Auto-Regressive (AR) ones. To this end, in this paper, we propose a novel dual diffusion-based framework namely MirrorDiff, to achieve semantic consistency with a symmetric image-to-text-to-image generation model, which acts like a mirror that maps the original input image into a regenerated image via the generated caption. Specifically, it first utilizes both pre-trained image encoder and text encoder to obtain image representation and textual representation respectively, then forwards the image representation and the noisy textual representation into a continuous diffusion model to output an intermediate sentence. To semantically align the intermediate sentence with the input image, a diffusion-based visual regenerator is employed to regenerate the input image conditioned on the intermediate sentence, resulting in a proposed visual regeneration loss. Different from most existing image captioning methods, MirrorDiff is a plug-and-play framework which can be plugged into many previous image captioning methods, and further evaluate the generated sentence via the visual similarity between the input image and the regenerated image. Extensive experiments on the MS COCO dataset show that our method achieves obvious improvements over state-of-the-art diffusion-based methods, up to 127.9 on CIDEr, and achieves competitive performance on multiple evaluation metrics over the auto-regressive methods trained on larger-scale datasets. Junbo Wang 0003, Liangyu Fu, Yining Zhu, Qiangguo Jin, Hongsong Wang 0001, Kun Hu 0008 |
ICMR | 5 |
| 2025 | DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image Captioning
Liangyu Fu, Junbo Wang 0003, Qiangguo Jin, Hongsong Wang 0001, Jing Ya, Linjiang Huang, Jiangbin Zheng 0001, Zhiyong Wang 0001 |
ACM Multimedia | 5 |
| 2025 | SoPo: Text-to-Motion Generation Using Semi-Online Preference OptimizationabstractText-to-motion generation is essential for advancing the creative industry but often presents challenges in producing consistent, realistic motions. To address this, we focus on fine-tuning text-to-motion models to consistently favor high-quality, human-preferred motions—a critical yet largely unexplored problem. In this work, we theoretically investigate the DPO under both online and offline settings, and reveal their respective limitation: overfitting in offline DPO, and biased sampling in online DPO. Building on our theoretical insights, we introduce Semi-online Preference Optimization (SoPo), a DPO-based method for training text-to-motion models using ``semi-online” data pair, consisting of unpreferred motion from online distribution and preferred motion in offline datasets. This method leverages both online and offline DPO, allowing each to compensate for the other’s limitations. Extensive experiments demonstrate that SoPo outperforms other preference alignment methods, with an MM-Dist of 3.25\% (vs e.g. 0.76\% of MoDiPO) on the MLD model, 2.91\% (vs e.g. 0.66\% of MoDiPO) on MDM model, respectively. Additionally, the MLD model fine-tuned by our SoPo surpasses the SoTA model in terms of R-precision and MM Dist. Visualization results also show the efficacy of our SoPo in preference alignment. Project page: https://xiaofeng-tan.github.io/projects/SoPo/. Xiaofeng Tan 0001, Hongsong Wang 0001, Xin Geng 0001, Pan Zhou 0002 |
NeurIPS | 2 |
| 2025 | Addressing Skewed Heterogeneity via Federated Prototype Rectification With PersonalizationabstractFederated learning (FL) is an efficient framework designed to facilitate collaborative model training across multiple distributed devices while preserving user data privacy. A significant challenge of FL is data-level heterogeneity, i.e., skewed or long-tailed distribution of private data. Although various methods have been proposed to address this challenge, most of them assume that the underlying global data are uniformly distributed across all clients. This article investigates data-level heterogeneity FL with a brief review and redefines a more practical and challenging setting called skewed heterogeneous FL (SHFL). Accordingly, we propose a novel federated prototype rectification with personalization (FedPRP) which consists of two parts: federated personalization and federated prototype rectification. The former aims to construct balanced decision boundaries between dominant and minority classes based on private data, while the latter exploits both interclass discrimination and intraclass consistency to rectify empirical prototypes. Experiments on three popular benchmarks show that the proposed approach outperforms current state-of-the-art methods and achieves balanced performance in both personalization and generalization. Shunxin Guo, Hongsong Wang 0001, Shuxia Lin, Zhiqiang Kou, Xin Geng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Robust 3D Face Alignment with Multi-Path Neural Architecture Searchabstract3D face alignment is a very challenging and fundamental problem in computer vision. Existing deep learning-based methods manually design different networks to regress either parameters of a 3D face model or 3D positions of face vertices. However, designing such networks relies on expert knowledge, and these methods often struggle to produce consistent results across various face poses. To address this limitation, we employ Neural Architecture Search (NAS) to automatically discover the optimal architecture for 3D face alignment. We propose a novel Multi-path One-shot Neural Architecture Search (MONAS) framework that leverages multi-scale features and contextual information to enhance face alignment across various poses. The MONAS comprises two key algorithms: Multi-path Networks Unbiased Sampling Based Training and Simulated Annealing based Multi-path One-shot Search. Experimental results on three popular benchmarks demonstrate the superior performance of the MONAS for both sparse alignment and dense alignment. Zhichao Jiang, Hongsong Wang 0001, Xi Teng, Baopu Li |
ICME | 2 |
| 2024 | Region-aware image-based human action retrieval with transformers
Hongsong Wang 0001, Jie Gui |
Comput. Vis. Image Underst. | 1 |
| 2024 | Dynamic heterogeneous federated learning with multi-level prototypes
Shunxin Guo, Hongsong Wang 0001, Xin Geng 0001 |
Pattern Recognit. | 2 |
| 2024 | Graph Convolution Based Efficient Re-Ranking for Visual RetrievalabstractVisual retrieval tasks such as image retrieval and person re-identification (Re-ID) aim at effectively and thoroughly searching images with similar content or the same identity. After obtaining retrieved examples, re-ranking is a widely adopted post-processing step to reorder and improve the initial retrieval results by making use of the contextual information from semantically neighboring samples. Prevailing re-ranking approaches update distance metrics and mostly rely on inefficient crosscheck set comparison operations while computing expanded neighbors based distances. In this work, we present an efficient re-ranking method which refines initial retrieval results by updating features. Specifically, we reformulate re-ranking based on Graph Convolution Networks (GCN) and propose a novel Graph Convolution based Re-ranking (GCR) for visual retrieval tasks via feature propagation. To accelerate computation for large-scale retrieval, a decentralized and synchronous feature propagation algorithm which supports parallel or distributed computing is introduced. In particular, the plain GCR is extended for cross-camera retrieval and an improved feature propagation formulation is presented to leverage affinity relationships across different cameras. It is also extended for video-based retrieval, and Graph Convolution based Re-ranking for Video (GCRV) is proposed by mathematically deriving a novel profile vector generation method for the tracklet. Without bells and whistles, the proposed approaches achieve state-of-the-art performances on seven benchmark datasets from three different tasks, i.e., image retrieval, person Re-ID and video-based person Re-ID. Yuqi Zhang 0001, Qi Qian 0001, Hongsong Wang 0001, Chong Liu 0002, Fan Wang 0019 |
IEEE Trans. Multim. | 3 |
| 2023 | Occluded Skeleton-Based Human Action Recognition with Dual Inhibition TrainingabstractRecently, skeleton-based human action recognition has received widespread attention in computer vision community. However, most existing research focuses on improving the recognition accuracy on complete skeleton data, while ignoring the performance on the incomplete skeleton data with occlusion or noise. This paper addresses occluded and noise-robust skeleton-based action recognition and presents a novel Dual Inhibition Training strategy. Specifically, we propose Part-aware and Dual-inhibition Graph Convolutional Network (PDGCN), which comprises of three parts: Input Skeleton Inhibition (ISI), Part-Aware Representation Learning (PARL) and Predicted Score Inhibition (PSI). The ISI and PSI are plug and play modules which could encourage the model to learn discriminative features from diversified body joints by effectively simulating key body part occlusions and random occlusions. The PARL module learns both the global and local representations from the whole body and body parts, respectively, and progressively fuses them during representation learning to enhance the model robustness under occlusions. Finally, we design different settings for occluded skeleton-based human action recognition to deep study this problem and better evaluate different approaches. Our approach achieves state-of-the-art results on different benchmarks and dramatically outperforms the recent skeleton-based action recognition approaches, especially under large-scale temporal occlusion. Zhenjie Chen, Hongsong Wang 0001, Jie Gui |
ACM Multimedia | 2 |
| 2023 | Learning Efficient Representations for Image-Based Patent Retrieval
Hongsong Wang 0001, Yuqi Zhang 0001 |
PRCV (7) | 1 |
| 2023 | Efficient Token-Guided Image-Text Retrieval With Consistent Multimodal Contrastive TrainingabstractImage-text retrieval is a central problem for understanding the semantic relationship between vision and language, and serves as the basis for various visual and language tasks. Most previous works either simply learn coarse-grained representations of the overall image and text, or elaborately establish the correspondence between image regions or pixels and text words. However, the close relations between coarse- and fine-grained representations for each modality are important for image-text retrieval but almost neglected. As a result, such previous works inevitably suffer from low retrieval accuracy or heavy computational cost. In this work, we address image-text retrieval from a novel perspective by combining coarse- and fine-grained representation learning into a unified framework. This framework is consistent with human cognition, as humans simultaneously pay attention to the entire sample and regional elements to understand the semantic content. To this end, a Token-Guided Dual Transformer (TGDT) architecture which consists of two homogeneous branches for image and text modalities, respectively, is proposed for image-text retrieval. The TGDT incorporates both coarse- and fine-grained retrievals into a unified framework and beneficially leverages the advantages of both retrieval approaches. A novel training objective called Consistent Multimodal Contrastive (CMC) loss is proposed accordingly to ensure the intra- and inter-modal semantic consistencies between images and texts in the common embedding space. Equipped with a two-stage inference method based on the mixed global and local cross-modal similarity, the proposed method achieves state-of-the-art retrieval performances with extremely low inference time when compared with representative recent approaches. Code is publicly available: github.com/LCFractal/TGDT. Chong Liu 0002, Yuqi Zhang 0001, Hongsong Wang 0001, Fan Wang 0019, Yan Huang 0008, Yidong Shen, Liang Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Velocity-to-velocity human motion forecasting
Hongsong Wang 0001, Liang Wang 0001, Jiashi Feng, Daquan Zhou |
Pattern Recognit. | 1 |
| 2021 | PVRED: A Position-Velocity Recurrent Encoder-Decoder for Human Motion PredictionabstractHuman motion prediction, which aims to predict future human poses given past poses, has recently seen increased interest. Many recent approaches are based on Recurrent Neural Networks (RNN) which model human poses with exponential maps. These approaches neglect the pose velocity as well as temporal relation of different poses, and tend to converge to the mean pose or fail to generate natural-looking poses. We therefore propose a novel Position-Velocity Recurrent Encoder-Decoder (PVRED) for human motion prediction, which makes full use of pose velocities and temporal positional information. A temporal position embedding method is presented and a Position-Velocity RNN (PVRNN) is proposed. We also emphasize the benefits of quaternion parameterization of poses and design a novel trainable Quaternion Transformation (QT) layer, which is combined with a robust loss function during training. We provide quantitative results for both short-term prediction in the future 0.5 seconds and long-term prediction in the future 0.5 to 1 seconds. Experiments on several benchmarks show that our approach considerably outperforms the state-of-the-art methods. In addition, qualitative visualizations in the future 4 seconds show that our approach could predict future human-like and meaningful poses in very long time horizons. Code is publicly available on GitHub: https://github.com/hongsong-wang/PVRNN. Hongsong Wang 0001, Jian Dong 0011, Bin Cheng 0001, Jiashi Feng |
IEEE Trans. Image Process. | 1 |
| 2021 | AFAN: Augmented Feature Alignment Network for Cross-Domain Object DetectionabstractUnsupervised domain adaptation for object detection is a challenging problem with many real-world applications. Unfortunately, it has received much less attention than supervised object detection. Models that try to address this task tend to suffer from a shortage of annotated training samples. Moreover, existing methods of feature alignments are not sufficient to learn domain-invariant representations. To address these limitations, we propose a novel augmented feature alignment network (AFAN) which integrates intermediate domain image generation and domain-adversarial training into a unified framework. An intermediate domain image generator is proposed to enhance feature alignments by domain-adversarial training with automatically generated soft domain labels. The synthetic intermediate domain images progressively bridge the domain divergence and augment the annotated source domain training data. A feature pyramid alignment is designed and the corresponding feature discriminator is used to align multi-scale convolutional features of different semantic levels. Last but not least, we introduce a region feature alignment and an instance discriminator to learn domain-invariant features for object proposals. Our approach significantly outperforms the state-of-the-art methods on standard benchmarks for both similar and dissimilar domain adaptations. Further extensive experiments verify the effectiveness of each component and demonstrate that the proposed network can learn domain-invariant representations. Hongsong Wang 0001, Shengcai Liao, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Learning content and style: Joint action recognition and person identification from human skeletons
Hongsong Wang 0001, Liang Wang 0001 |
Pattern Recognit. | 1 |
| 2018 | Cross-Agent Action RecognitionabstractAn action is something which is done by an agent. Most action recognition researchers merely focus on the actions to be recognized, and ignore the differences of agents. Philosophers and behaviorists discover that actions are common among many species, but are performed in different ways and with different levels of sophistication. In this paper, in order to bridge action recognition tasks between different agents, we introduce a new problem, cross-agent action recognition, i.e., recognizing action for one particular agent (target) while training from other agents (source). We model this problem under three different scenarios: single source and single target, multiple sources and single target, and multiple sources and multiple targets. To this end, corresponding methods based on transfer learning are proposed to address these problems. We further design three different strategies to model the situation when a partial labeled data is provided for the target. Experimental results show that the performances of the transfer method are generally better than those of the comparative method without transfer learning, especially when we have multiple sources. Particularly, the transfer method outperforms the others significantly when the source is a human adult. In addition, cross-agent method significantly improves the results when partially labeled data is provided for the target. These demonstrate that for action recognition, knowledge can be transferred across different agents. A straightforward application of this finding is to use human action (training data is abundant) data to enhance animal action recognition. Hongsong Wang 0001, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Beyond Joints: Learning Representations From Primitive Geometries for Skeleton-Based Action Recognition and DetectionabstractRecently, skeleton-based action recognition becomes popular owing to the development of cost-effective depth sensors and fast pose estimation algorithms. Traditional methods based on pose descriptors often fail on large-scale datasets due to the limited representation of engineered features. Recent recurrent neural networks (RNN) based approaches mostly focus on the temporal evolution of body joints and neglect the geometric relations. In this paper, we aim to leverage the geometric relations among joints for action recognition. We introduce three primitive geometries: joints, edges, and surfaces. Accordingly, a generic end-to-end RNN based network is designed to accommodate the three inputs. For action recognition, a novel viewpoint transformation layer and temporal dropout layers are utilized in the RNN based network to learn robust representations. And for action detection, we first perform frame-wise action classification, then exploit a novel multi-scale sliding window algorithm. Experiments on the large-scale 3D action recognition benchmark datasets show that joints, edges, and surfaces are effective and complementary for different actions. Our approaches dramatically outperform the existing state-of-the-art methods for both tasks of action recognition and action detection. Hongsong Wang 0001, Liang Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Modeling Temporal Dynamics and Spatial Configurations of Actions Using Two-Stream Recurrent Neural NetworksabstractRecently, skeleton based action recognition gains more popularity due to cost-effective depth sensors coupled with real-time skeleton estimation algorithms. Traditional approaches based on handcrafted features are limited to represent the complexity of motion patterns. Recent methods that use Recurrent Neural Networks (RNN) to handle raw skeletons only focus on the contextual dependency in the temporal domain and neglect the spatial configurations of articulated skeletons. In this paper, we propose a novel two-stream RNN architecture to model both temporal dynamics and spatial configurations for skeleton based action recognition. We explore two different structures for the temporal stream: stacked RNN and hierarchical RNN. Hierarchical RNN is designed according to human body kinematics. We also propose two effective methods to model the spatial structure by converting the spatial graph into a sequence of joints. To improve generalization of our model, we further exploit 3D transformation based data augmentation techniques including rotation and scaling transformation to transform the 3D coordinates of skeletons during training. Experiments on 3D action recognition benchmark datasets show that our method brings a considerable improvement for a variety of actions, i.e., generic actions, interaction activities and gestures. Hongsong Wang 0001, Liang Wang 0001 |
CVPR | 1 |
| 2016 | How scenes imply actions in realistic videos?abstractPeople drive on the road and eat in the kitchen. Can the road imply driving or the kitchen imply eating? This paper addresses such a problem by studying the relations between actions and scenes. To get effective scene representation, we use a deep convolutional neural networks (CNN) model trained from a scene-centric database to predict scene responses for videos. We employ two encoding schemes based on frame features to represent the scene and its changes, respectively. We conduct experiments on two challenging datasets, HMDB51 and Hollywood2, and compare action recognition results of different encodings based on different scene features. Our results demonstrate that scene features, when combined with motion features, improve the state-of-the-art results for action recognition. Finally, we explore the relationship between actions and scenes by analyzing scene preferences to a particular action qualitatively and quantitatively. Hongsong Wang 0001, Wei Wang 0115, Liang Wang 0001 |
ICIP | 1 |