Hong Liu 0008

dblp:29/5010-8 · DBLP profile ↗
← Back
229ranked-venue papers
75as first author
81since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 137 · 42 first-author · 53 since 2021Artificial intelligence and machine learning · 114 · 31 first-author · 45 since 2021Systems, architecture and hardware · 25 · 17 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 13 first-authorHuman-computer interaction and ubiquitous computing · 15 · 11 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
abstract
Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We propose MaskClu, a novel unsupervised pre-training method for ViTs on 3D point clouds that integrates masked point modeling with clustering-based learning. MaskClu is designed to reconstruct both cluster assignments and cluster centers from masked point clouds, thus encouraging the model to capture dense semantic information. Additionally, we introduce a global contrastive learning mechanism that enhances instance-level feature learning by contrasting different masked views of the same point cloud. By jointly optimizing these complementary objectives, i.e., dense semantic reconstruction, and instance-level contrastive learning. MaskClu enables ViTs to learn richer and more semantically meaningful representations from 3D point clouds. We validate the effectiveness of MaskClu via multiple 3D tasks, including part segmentation, semantic segmentation, object detection, and classification, setting new competitive results.
Bin Ren 0005, Xiaoshui Huang, Mengyuan Liu 0001, Hong Liu 0008, Fabio Poiesi, Nicu Sebe, Guofeng Mei
AAAI4
2026 Debiased Multiplex Tokenizer for Efficient Map-Free Visual Relocalization
abstract
Image-based feature representation plays a critical role in visual localization, enabling robots to estimate their position and orientation in GPS-denied environments. However, this task is often undermined by significant variations in camera viewpoints and scene appearances. Recently, map-free visual relocalization (MFVR) has emerged as a promising paradigm due to its compatibility with lightweight deployment and privacy isolation on mobile devices. In this paper, we propose the Debiased Multiplex Tokenizer (DeMT) as a novel method for versatile and efficient MFVR. Specifically, DeMT performs relative pose regression through an integrated framework built upon a pretrained vision Mamba encoder, comprising three key modules: First, Multiplex Interactive Tokenization yields robust image tokens with non-local affinities and cross-domain descriptions; Second, Debiased Anchor Registration facilitates anchor token matching through proximity graph retrieval and causal pointer attribution; Third, Geometry-Informed Pose Regression empowers multi-layer perceptrons with a gating mechanism and spectral normalization to support both pair-wise and multi-view modes. Extensive evaluations across nine public datasets demonstrate that DeMT substantially outperforms existing baselines and ablation variants in diverse indoor and outdoor environments.
Hong Liu 0008, Shengquan Li 0001, Peifeng Jiang, Runwei Ding
AAAI2
2026 H2OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
abstract
Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper, we present a hierarchical plug-and-play pruning-and-recovering framework, calledHierarchicalHourglassTokenizer (H2OT), for efficient transformer-based 3D human pose estimation from videos. H2OT begins with progressively pruning pose tokens of redundant frames and ends with recovering full-length sequences, resulting in a few pose tokens in the intermediate transformer blocks and thus improving the model efficiency. It works with two key modules, namely, a Token Pruning Module (TPM) and a Token Recovering Module (TRM). TPM dynamically selects a few representative tokens to eliminate the redundancy of video frames, while TRM restores the detailed spatio-temporal information based on the selected tokens, thereby expanding the network output to the original full-length temporal resolution for fast inference. Our method is general-purpose: it can be easily incorporated into common VPT models on bothseq2seqandseq2framepipelines while effectively accommodating different token pruning and recovery strategies. In addition, our H2OT reveals that maintaining the full pose sequence is unnecessary, and a few pose tokens of representative frames can achieve both high efficiency and estimation accuracy. Extensive experiments on multiple benchmark datasets demonstrate both the effectiveness and efficiency of the proposed method. Code and models are available athttps://github.com/NationalGAILab/HoT.
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Pichao Wang, Shijian Lu, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Global-local distillation network-based audio-visual speaker tracking with incomplete modalities
Yidi Li 0001, Zhenhuan Xu, Weiwei Wan, Hong Liu 0008
Pattern Recognit.5
2026 MXPose: Multiplex interactive learning for multi-view 3D Human Pose Estimation
Wanruo Zhang, Zhou Guan, Mengyuan Liu 0001, Bin Ren 0005, Hong Liu 0008
Pattern Recognit.5
2026 DSVTformer: Dual-stream spatial-view-temporal transformer for multi-view 3D human pose estimation
Wanruo Zhang, Wenhao Li 0002, Hong Liu 0008
Pattern Recognit.4
2026 Uncertainty-Aware Testing-Time Optimization for 3D Human Pose Estimation
abstract
Although data-driven methods have achieved success in 3D human pose estimation, they often suffer from domain gaps and exhibit limited generalization. In contrast, optimization-based methods excel in fine-tuning for specific cases but are generally inferior to data-driven methods in overall performance. We observe that previous optimization-based methods commonly rely on projection constraint, which only ensures alignment in 2D space, potentially leading to the overfitting problem. To address this, we propose an Uncertainty-Aware testing-time Optimization (UAO) framework, which keeps the prior information of pre-trained model and alleviates the overfitting problem using the uncertainty of joints. Specifically, during the training phase, we design an effective 2D-to-3D network for estimating the corresponding 3D pose while quantifying the uncertainty of each 3D joint. For optimization during testing, the proposed optimization framework freezes the pre-trained model and optimizes only a latent state. Projection loss is then employed to ensure the generated poses are well aligned in 2D space for high-quality optimization. Furthermore, we utilize the uncertainty of each joint to determine how much each joint is allowed for optimization. The effectiveness and superiority of the proposed framework are validated through extensive experiments on challenging datasets: Human3.6M, MPI-INF-3DHP, and 3DPW. Notably, our approach outperforms the previous best result by a large margin of 5.5% on Human3.6M.
Ti Wang, Mengyuan Liu 0001, Hong Liu 0008, Bin Ren 0005, Yingxuan You, Wenhao Li 0002, Nicu Sebe, Xia Li 0005
IEEE Trans. Multim.3
2025 TCPFormer: Learning Temporal Correlation with Implicit Pose Proxy for 3D Human Pose Estimation
abstract
Recent multi-frame lifting methods have dominated the 3D human pose estimation. However, previous methods ignore the intricate dependence within the 2D pose sequence and learn single temporal correlation. To alleviate this limitation, we propose TCPFormer, which leverages an implicit pose proxy as an intermediate representation. Each proxy within the implicit pose proxy can build one temporal correlation therefore helping us learn a more comprehensive temporal correlation of human motion. Specifically, our method consists of three key components: Proxy Update Module (PUM), Proxy Invocation Module (PIM), and Proxy Attention Module (PAM). PUM first uses pose features to update the implicit pose proxy, enabling it to store representative information from the pose sequence. PIM then invocates and integrates the pose proxy with the pose sequence to enhance the motion semantics of each pose. Finally, PAM leverages the above mapping between the pose sequence and pose proxy to enhance the temporal correlation of the whole pose sequence. Experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that our proposed TCPFormer outperforms the previous state-of-the-art methods.
Hong Liu 0008, Wenhao Li 0002
AAAI3
2025 PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes
abstract
The integration of RGB and depth modalities significantly enhances the accuracy of segmenting complex indoor scenes, with depth data from RGB-D cameras playing a crucial role in this improvement. However, collecting an RGB-D dataset is more expensive than an RGB dataset due to the need for specialized depth sensors. Aligning depth and RGB images also poses challenges due to sensor positioning and issues like missing data and noise. In contrast, Pseudo Depth (PD) from high-precision depth estimation algorithms can eliminate the dependence on RGB-D sensors and alignment processes, as well as provide effective depth information and show significant potential in semantic segmentation. Therefore, to explore the practicality of utilizing pseudo depth instead of real depth for semantic segmentation, we design an RGB-PD segmentation pipeline to integrate RGB and pseudo depth and propose a Pseudo Depth Aggregation Module (PDAM) for fully exploiting the informative clues provided by the diverse pseudo depth maps. The PDAM aggregates multiple pseudo depth maps into a single modality, making it easily adaptable to other RGB-D segmentation methods. In addition, the pre-trained diffusion model serves as a strong feature extractor for RGB segmentation tasks, but multi-modal diffusion-based segmentation methods remain unexplored. Therefore, we present a Pseudo Depth Diffusion Model (PDDM) that adopts a large-scale text-image diffusion model as a feature extractor and a simple yet effective fusion strategy to integrate pseudo depth. To verify the applicability of pseudo depth and our PDDM, we perform extensive experiments on the NYUv2 and SUNRGB-D datasets. The experimental results demonstrate that pseudo depth can effectively enhance segmentation performance, and our PDDM achieves state-of-the-art performance, outperforming other methods by +6.98 mIoU on NYUv2 and +2.11 mIoU on SUNRGB-D.
Xinhua Xu, Hong Liu 0008, Jianbing Wu
AAAI2
2025 SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose Estimation
abstract
Recently, transformer-based methods have been introduced to estimate 3D human pose from multiple views by aggregating the spatial-temporal information of human joints to achieve the lifting of 2D to 3D. However, previous approaches cannot model the inter-frame correspondence of each view's joint individually, nor can they directly consider all view interactions at each time, leading to insufficient learning of multi-view associations. To address this issue, we propose a Spatial-View-Temporal transformer (SVTformer) to decouple spatial-view-temporal information in sequential order for correlation learning and model dependencies between them in a local-to-global manner. SVTformer includes an attended Spatial-View-Temporal (SVT) patch embedding to attentively capture the local features of the input poses and stacked SVT encoders to extract global spatial-view-temporal dependencies progressively. Specifically, SVT encoders perform three reconstructions sequentially to attended features with the learning through view decoupling for temporal-enhanced spatial correlation, temporal decoupling for spatial-enhanced view correlation, and another view decoupling for spatial-enhanced temporal relationship. This decoupling-coupling-decoupling multi-view scheme enables us to alternatively model the inter-joint spatial relationships, cross-view dependencies, and temporal motion associations. We evaluate the proposed SVTformer on three popular 3D HPE datasets, and it yields state-of-the-art performance. It effectively deals with ill-posed problems and enhances the accuracy of 3D human pose estimation.
Wanruo Zhang, Hong Liu 0008, Wenhao Li 0002
AAAI3
2025 UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
abstract
Point cloud videos capture dynamic 3D motion while reducing the effects of lighting and viewpoint variations, making them highly effective for recognizing subtle and continuous human actions. Although Selective State Space Models (SSMs) have shown good performance in sequence modeling with linear complexity, the spatio-temporal disorder of point cloud videos hinders their unidirectional modeling when directly unfolding the point cloud video into a 1D sequence through temporally sequential scanning. To address this challenge, we propose the Unified Spatio-Temporal State Space Model (UST-SSM), which extends the latest advancements in SSMs to point cloud videos. Specifically, we introduce Spatial-Temporal Selection Scanning (STSS), which reorganizes unordered points into semantic-aware sequences through prompt-guided clustering, thereby enabling the effective utilization of points that are spatially and temporally distant yet similar within the sequence. For missing 4D geometric and motion details, Spatio-Temporal Structure Aggregation (STSA) aggregates spatio-temporal features and compensates. To improve temporal interaction within the sampled sequence, Temporal Interaction Sampling (TIS) enhances fine-grained temporal dependencies through non-anchor frame utilization and expanded receptive fields. Experimental results on the MSR-Action3D, NTU RGB+D, and Synthia 4D datasets validate the effectiveness of our method. Our code is available at https://github.com/wangzy01/UST-SSM.
Peiming Li, Yulin Yuan, Hong Liu 0008, Xiangming Meng, Junsong Yuan 0001, Mengyuan Liu 0004
ICCV4
2025 Recognizing Actions From Robotic View for Natural Human-Robot Interaction
Peiming Li, Hong Liu 0008, Zhichao Deng, Can Wang 0006, Jun Liu 0036, Junsong Yuan 0001, Mengyuan Liu 0001
ICCV3
2025 A Unified End-to-End Network for Category-Level and Instance-Level Object Pose Estimation from RGB Images
abstract
Accurately estimating the 6-DoF pose of objects is a fundamental challenge in computer vision and robotics. While category-level pose estimation based on RGBD data has achieved good performance in recent years, estimating poses solely from RGB images remains a significant challenge. Existing RGB-based category-level methods primarily focus on recovering object point clouds from RGB images, and pose prediction is not performed end-to-end by a network. This paper presents a Category-level and Instance-level Pose Estimation Network (CIPE), which models pose estimation as a set prediction problem and enables direct pose regression from RGB images. To further enhance the network's ability to learn object poses, first, a novel learnable rotation representation that redefines rotation learning within Euclidean space is introduced to facilitate rotation regression. Additionally, we propose a prior-query fusion strategy that utilizes a pre-trained point cloud feature extraction network to integrate categorical object features with bounding boxes, thereby improving the incorporation of category information. Experimental results demonstrate that CIPE significantly outperforms existing RGB-based methods on both category-level and instance-level datasets. The code is available at https://github.com/jialeren/CIPE.
Jiale Ren, Hong Liu 0008, Peifeng Jiang
ICRA2
2025 PRIDEV: A Plug-and-Play Refinement for Improved Depth Estimation in Videos
abstract
Monocular video depth estimation is a key challenge in computer vision, highlighting its importance in visual understanding. Monocular depth estimation models trained on single images achieve impressive results on individual frames but often lack temporal consistency when applied to videos, leading to flickering and artifacts. Current video depth estimation methods often rely on additional optical flow or camera poses, which are limited by their accuracy, complex design, and lack robustness. Specially, we propose a plug-and-play method that seamlessly transfers the robustness of image depth estimation to video depth estimation. By leveraging powerful priors from image depth estimation, our method enhances the performance of video depth estimation without requiring additional conditional inputs or extensive pretraining on large and expensive video datasets. We introduce the Temporal Depth Stabilization Module (TDSM), which can seamlessly inflate an image monocular depth estimation model into a video depth estimation model, enabling unified modeling of depth across video sequences and capturing the temporal cues in video. We validate the effectiveness and efficiency of our method across various datasets (e.g., normal and challenging conditions) and different backbones. Extensive experiments demonstrate that our simple and effective method significantly improves monocular depth estimation networks, achieving new state-of-the-art accuracy in both spatial and temporal dimensions.
Hong Liu 0008, Jianbing Wu, Xinhua Xu
ICRA2
2025 TCNet: A Temporally Consistent Network for Self-supervised Monocular Depth Estimation
abstract
Despite significant advances in self-supervised monocular depth estimation methods, achieving temporally consistent and accurate depth maps from frame sequences remains a formidable challenge. Existing approaches often estimate depth maps for individual frames in isolation, neglecting the rich geometric and temporal coherence present across frames. Consequently, this oversight leads to temporally inconsistent outputs, resulting in noticeable temporal flickering artifacts. In response, this paper presents TCNet, a Temporal Consistent Network for self-supervised monocular depth estimation. Specifically, we propose an Inter-frame Temporal Fusion (ITF) module to emphasize the influence of preceding images on the depth estimation of the current frame. The Temporal Consistency Loss (TCL) is proposed to leverage the temporal constraints between the depth maps of adjacent frames. Besides, TCNet can also be applied to both single-frame and multi-frame scenarios during inference. Experimental evaluations on the KITTI dataset demonstrate that our method surpasses state-of-the-art depth estimation methods in accuracy and temporal consistency. Our code will be made public.
Hong Liu 0008, Jianbing Wu
IROS2
2025 CLIP-6D: Empowering CLIP as a Zero-Shot 6D Pose Estimator Through Generalizable Object-Specific Representations
abstract
The open-vocabulary paradigm enhances 6D object pose estimation by leveraging language cues to transfer learned representations from seen to unseen objects, yet its performance suffers from a misalignment between vision-language representations and the 6D pose space. This challenge is compounded by the intrinsic lack of pose awareness in models like CLIP (Contrastive Language-Image Pre-Training). To address this limitation, we introduce CLIP-6D, which equips CLIP with the ability to estimate poses through the learning of generalizable representations specific to objects. CLIP-6D consists of three key components: (1) an innovative self-alignment strategy that enables CLIP to derive geometric representations from RGB images by leveraging its inherent feature extractor; % utilizing its inherent feature extraction capability; (2) a multiplex representation interactive learning method efficiently bridges the heterogeneous representations of object category priors, geometry, and spatial correlations; (3) a lightweight adapter using knowledge distillation improves CLIP's capture of detailed semantic representations. Experiments show that CLIP-6D achieves an improvement of 10.8% and 16.4% in the metric 5° 2 cm and 10° 5 cm for zero-shot generalization and achieves a speedup of 5.4 FPS in inference over current state-of-the-art methods. The source code and models are available at https://github.com/whoawong/CLIP-6D.git.
Hong Liu 0008, Jiale Ren, Mingxin Tan, Zhongzien Jiang
ACM Multimedia2
2025 PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
abstract
The integration of point and voxel representations is becoming more common in Light Detection and Ranging (LiDAR)-based 3D object detection. However, existing fusion strategies suffer from ineffective semantic alignment and contextual information loss, while relying solely on point features within regions of interest leads to geometric detail degradation and limited local–global feature integration. To tackle these challenges, we propose the Point-Voxel Attention Fusion Network (PVAFN), a novel two-stage 3D object detector that introduces a point-voxel attention fusion module based on dual-gated cross-modal interaction and a multi-pooling strategy based on density-space awareness. During the feature extraction and fusion stage, a dual-gated hierarchical attention mechanism is proposed to dynamically fuse three heterogeneous modalities—keypoint-based geometric details, voxel-wise local regularity, and Bird’s-Eye-View (BEV)-level global semantics—through learnable gating functions. In the refinement stage, a density-spatial-aware multi-pooling enhancement module is designed to synergize density-aware cluster pooling and multi-scale spatial-aware pyramid pooling, efficiently capturing key geometric details and fine-grained shape structures. This design enhances the integration of local and global features while enabling adaptive multi-scale context modeling and spatially sensitive feature aggregation. Extensive experiments on the KITTI and Waymo benchmark datasets demonstrate that PVAFN achieves promising detection accuracy in 3D mean Average Precision. • PVAFN: A novel network for 3D object detection, fusing keypoint, voxel, and BEV features dynamically. • Stage-I: Dual-gated hierarchical attention fusion for bidirectional point-voxel feature calibration. • Stage-II: Density-spatial-aware multi-pooling module enhances local and global geometric perception. • SOTA Performance: PVAFN achieves highest AP on KITTI and Waymo benchmarks for autonomous driving.
Yidi Li 0001, Bin Ren 0005, Wenhao Li 0002, Hong Liu 0008, Nicu Sebe
Expert Syst. Appl.7
2025 GraphMLP: A graph MLP-like architecture for 3D human pose estimation
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Tianyu Guo 0001, Ti Wang, Hao Tang 0005, Nicu Sebe
Pattern Recognit.3
2025 Frequency-Aware Self-Supervised Group Activity Recognition with skeleton sequences
Mengyuan Liu 0001, Hong Liu 0008, Jinyan Zhang, Peini Guo, Ruijia Fan
Pattern Recognit.3
2025 Dual Attention Guidance Network for Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation shows great promise since only a single camera is required. However, most existing methods fail to model the geometric structure of objects, leading to poor performance in object boundary depth estimation. To overcome these shortcomings, a dual attention guidance network (DAG-Net), containing two complementary modules termed depth-guided attention module (DAM) and semantic-guided multi-modal attention module (SAM), is proposed in this paper. The DAM utilizes depth features to guide semantic features through multi-head attention. When semantic features are well learned, they guide depth features to learn useful geometric representations through backpropagation. Besides, the SAM is proposed to incorporate multi-modal data from depth estimation and semantic segmentation predictions at different scales. To eliminate the mutual interference between DAM and SAM, we also propose a two-stage training strategy to adjust the convergence direction during the training process. The effectiveness of our proposed DAG-Net is qualitatively and quantitatively verified by various experiments on KITTI, Cityscapes, and Make3D datasets, showing outstanding performance compared with the state-of-the-art methods.
Hong Liu 0008, Guoliang Hua, Hao Tang 0005, Yidi Li 0001, Weibo Huang
IEEE Trans. Circuits Syst. Video Technol.2
2025 Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction Recognition
abstract
Recognizing interactive actions, including hand-to-hand interaction and human-to-human interaction, has attracted increasing attention for various applications in the field of video analysis and human–robot interaction. Considering the success of graph convolution in modeling topology-aware features from skeleton data, recent methods commonly operate graph convolution on separate entities and use late fusion for interactive action recognition, which can barely model the mutual semantic relationships between pairwise entities. To this end, we propose a mutual excitation graph convolutional network (me-GCN) by stacking mutual excitation graph convolution (me-GC) layers. Specifically, me-GC uses a mutual topology excitation module to firstly extract adjacency matrices from individual entities and then adaptively model the mutual constraints between them. Moreover, me-GC extends the above idea and further uses a mutual feature excitation module to extract and merge deep features from pairwise entities. Compared with graph convolution, our proposed me-GC gradually learns mutual information in each layer and each stage of graph convolution operations. Extensive experiments on a challenging hand-to-hand interaction dataset, i.e., the Assembely101 dataset, and two large-scale human-to-human interaction datasets, i.e., NTU60-Interaction and NTU120-Interaction consistently verify the superiority of our proposed method, which outperforms the state-of-the-art GCN-based and Transformer-based methods.
Mengyuan Liu 0001, Chen Chen 0015, Songtao Wu, Fanyang Meng, Hong Liu 0008
IEEE Trans. Hum. Mach. Syst.5
2025 NanoHTNet: Nano Human Topology Network for Efficient 3D Human Pose Estimation
abstract
The widespread application of 3D human pose estimation (HPE) is limited by resource-constrained edge devices like Jetson Nano, requiring more efficient models. A key approach to enhancing efficiency involves designing networks based on the structural characteristics of input data. However, effectively utilizing the structural priors in human skeletal inputs remains challenging. To address this, we leverage both explicit and implicit spatio-temporal priors of the human body through innovative model design and a pre-training proxy task. First, we propose a Nano Human Topology Network (NanoHTNet), a tiny 3D HPE network with stacked Hierarchical Mixers to capture explicit features. Specifically, the spatial Hierarchical Mixer efficiently learns the human physical topology across multiple semantic levels, while the temporal Hierarchical Mixer with discrete cosine transform and low-pass filtering captures local instantaneous movements and global action coherence. Moreover, Efficient Temporal-Spatial Tokenization (ETST) is introduced to enhance spatio-temporal interaction and reduce computational complexity significantly. Second, PoseCLR is proposed as a general pre-training method based on contrastive learning for 3D HPE, aimed at extracting implicit representations of human topology. By aligning 2D poses from diverse viewpoints in the proxy task, PoseCLR aids 3D HPE encoders like NanoHTNet in more effectively capturing the high-dimensional features of the human body, leading to further performance improvements. Extensive experiments verify that NanoHTNet with PoseCLR outperforms other state-of-the-art methods in efficiency, making it ideal for deployment on edge devices like the Jetson Nano. Code and models are available at https://github.com/vefalun/NanoHTNet.
Jialun Cai, Mengyuan Liu 0001, Hong Liu 0008, Shuheng Zhou 0001, Wenhao Li 0002
IEEE Trans. Image Process.3
2025 HYRE: Hybrid Regressor for 3D Human Pose and Shape Estimation
abstract
Regression-based 3D human pose and shape estimation often fall into one of two different paradigms. Parametric approaches, which regress the parameters of a human body model, tend to produce physically plausible but image-mesh misalignment results. In contrast, non-parametric approaches directly regress human mesh vertices, resulting in pixel-aligned but unreasonable predictions. In this paper, we consider these two paradigms together for a better overall estimation. To this end, we propose a novel HYbrid REgressor (HYRE) that greatly benefits from the joint learning of both paradigms. The core of our HYRE is a hybrid intermediary across paradigms that provides complementary clues to each paradigm at the shared feature level and fuses their results at the part-based decision level, thereby bridging the gap between the two. We demonstrate the effectiveness of the proposed method through both quantitative and qualitative experimental analyses, resulting in improvements for each approach and ultimately leading to better hybrid results. Our experiments show that HYRE outperforms previous methods on challenging 3D human pose and shape benchmarks.
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Bin Ren 0005, Xia Li 0005, Yingxuan You, Nicu Sebe
IEEE Trans. Image Process.3
2025 MiLNet: Multiplex Interactive Learning Network for RGB-T Semantic Segmentation
abstract
Semantic segmentation methods enhance robust and reliable understanding under adverse illumination conditions by integrating complementary information from visible and thermal infrared (RGB-T) images. Existing methods primarily focus on designing various feature fusion modules between different modalities, overlooking that feature learning is the critical aspect of scene understanding. In this paper, we propose a novel module-free Multiplex Interactive Learning Network (MiLNet) for RGB-T semantic segmentation, which adeptly integrates multi-model, multi-modal, and multi-level feature learning, fully exploiting the potential of multiplex feature interaction. Specifically, robust knowledge is transferred from the vision foundation model to our task-specific model to enhance its segmentation performance. In the task-specific model, an asymmetric simulated learning strategy is introduced to facilitate mutual learning of geometric and semantic information between high- and low-level features across modalities. Additionally, an inverse hierarchical fusion strategy based on feature learning pairs is adopted and further refined using multilabel and multiscale supervision. Experimental results on the MFNet and PST900 datasets demonstrate that MiLNet outperforms state-of-the-art methods in terms of mIoU. As a limitation, the model's performance under few-sample conditions could be improved further. The code and results of our method are available at https://github.com/Jinfu-pku/MiLNet.
Hong Liu 0008, Xia Li 0005, Jiale Ren, Xinhua Xu
IEEE Trans. Image Process.2
2025 STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking
abstract
Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Recently, several fusion methods have been proposed to model the correlation in multiple modalities. However, for the speaker tracking problem, the cross-modal interaction between audio and visual signals hasn't been well exploited. To this end, we present a novel Speaker Tracking Network (STNet) with a deep audio-visual fusion model in this work. We design a visual-guided acoustic measurement method to fuse heterogeneous cues in a unified localization space, which employs visual observations via a camera model to construct the enhanced acoustic map. For feature fusion, a cross-modal attention module is adopted to jointly model multi-modal contexts and interactions. The correlated information between audio and visual features is further interacted in the fusion model. Moreover, the STNet-based tracker is applied to multi-speaker cases by a quality-aware module, which evaluates the reliability of multi-modal observations to achieve robust tracking in complex scenarios. Experiments on the AV16.3 and CAV3D datasets show that the proposed STNet-based tracker outperforms uni-modal methods and state-of-the-art audio-visual speaker trackers.
Yidi Li 0001, Hong Liu 0008, Bing Yang 0004
IEEE Trans. Multim.2
2024 Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation
abstract
Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper, we present a plug-and-play pruning-and-recovering framework, called Hourglass Tokenizer (HoT), for efficient transformer-based 3D human pose estimation from videos. Our HoT begins with pruning pose tokens of re-dundant frames and ends with recovering full-length tokens, resulting in a few pose tokens in the intermediate transformer blocks and thus improving the model efficiency. To effectively achieve this, we propose a token pruning cluster (TPC) that dynamically selects a few representative tokens with high semantic diversity while eliminating the redundancy of video frames. In addition, we develop a token recovering attention (TRA) to restore the detailed spatio-temporal information based on the selected tokens, thereby expanding the network output to the original full-length temporal resolution for fast inference. Extensive experiments on two benchmark datasets (i.e., Human3.6M and MPI-INF-3DHP) demonstrate that our method can achieve both high efficiency and estimation accuracy compared to the original VPT models. For instance, applying to MotionBERT and MixSTE on Hu-man3.6M, our HoT can save nearly 50% FLOPs without sacrificing accuracy and nearly 40% FLOPs with only 0.2% accuracy drop, respectively. Code and models are available at https://github.com/NationalGAILab/HoT.
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Pichao Wang, Jialun Cai, Nicu Sebe
CVPR3
2024 AttA-NET: Attention Aggregation Network for Audio-Visual Emotion Recognition
abstract
In video-based emotion recognition, effective multi-modal fusion techniques are essential to leverage the complementary relationship between audio and visual modalities. Recent attention-based fusion methods are widely leveraged for capturing modal-shared properties. However, they often ignore the modal-specific properties of audio and visual modalities and the unalignment of model-shared emotional semantic features. In this paper, an Attention Aggregation Network (AttA-NET) is proposed to address these challenges. An attention aggregation module is proposed to get modal-shared properties effectively. This module comprises similarity-aware enhancement blocks and a contrastive loss that facilitates aligning audio and visual semantic features. Moreover, an auxiliary uni-modal classifier is introduced to obtain modal-specific properties, in which intra-modal discriminative features are fully extracted. Under joint optimization of uni-modal and multi-modal classification loss, modal-specific information can be infused. Extensive experiments on RAVDESS and PKU-ER datasets validate the superiority of AttA-NET. The code is available at: https://github.com/NariFan2002/AttA-NET.
Ruijia Fan, Hong Liu 0008, Yidi Li 0001, Peini Guo, Ti Wang
ICASSP2
2024 Dual-Branch Graph Transformer Network for 3D Human Mesh Reconstruction from Video
abstract
Human Mesh Reconstruction (HMR) from monocular video plays an important role in human-robot interaction and collaboration. However, existing video-based human mesh reconstruction methods face a trade-off between accurate reconstruction and smooth motion. These methods design networks based on either RNNs or attention mechanisms to extract local temporal correlations or global temporal dependencies, but the lack of complementary long-term information and local details limits their performance. To address this problem, we propose a Dual-branch Graph Transformer network for 3D human mesh Reconstruction from video, named DGTR. DGTR employs a dual-branch network including a Global Motion Attention (GMA) branch and a Local Details Refine (LDR) branch to par-allelly extract long-term dependencies and local crucial information, helping model global human motion and local human details (e.g., local motion, tiny movement). Specifically, GMA utilizes a global transformer to model long-term human motion. LDR combines modulated graph convolutional networks and the transformer framework to aggregate local information in adjacent frames and extract crucial information of human details. Experiments demonstrate that our DGTR outperforms state-of-the-art video-based methods in reconstruction accuracy and maintains competitive motion smoothness. Moreover, DGTR utilizes fewer parameters and FLOPs, which validate the effectiveness and efficiency of the proposed DGTR. Code is publicly available at https://github.com/TangTao-PKU/DGTR.
Hong Liu 0008, Yingxuan You, Ti Wang, Wenhao Li 0002
IROS2
2024 ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion Models
abstract
Grasp generation aims to create complex hand-object interactions with a specified object. While traditional approaches for hand generation have primarily focused on visibility and diversity under scene constraints, they tend to overlook the fine-grained hand-object interactions such as contacts, resulting in inaccurate and undesired grasps. To address these challenges, we propose a controllable grasp generation task and introduce ClickDiff, a controllable conditional generation model that leverages a fine-grained Semantic Contact Map (SCM). Particularly when synthesizing interactive grasps, the method enables the precise control of grasp synthesis through either user-specified or algorithmically predicted Semantic Contact Map. Specifically, to optimally utilize contact supervision constraints and to accurately model the complex physical structure of hands, we propose a Dual Generation Framework. Within this framework, the Semantic Conditional Module generates reasonable contact maps based on fine-grained contact information, while the Contact Conditional Module utilizes contact maps alongside object point clouds to generate realistic grasps. We evaluate the evaluation criteria applicable to controllable grasp generation. Both unimanual and bimanual generation experiments on GRAB and ARCTIC datasets verify the validity of our proposed method, demonstrating the efficacy and robustness of ClickDiff, even with previously unseen objects. Our code is available at https://github.com/adventurer-w/ClickDiff.
Peiming Li, Mengyuan Liu 0004, Hong Liu 0008, Chen Chen 0001
ACM Multimedia4
2024 ARTS: Semi-Analytical Regressor using Disentangled Skeletal Representations for Human Mesh Recovery from Videos
abstract
Although existing video-based 3D human mesh recovery methods have made significant progress, simultaneously estimating human pose and shape from low-resolution image features limits their performance. These image features lack sufficient spatial information about the human body and contain various noises (e.g., background, lighting, and clothing), which often results in inaccurate pose and inconsistent motion. Inspired by the rapid advance in human pose estimation, we discover that compared to image features, skeletons inherently contain accurate human pose and motion. Therefore, we propose a novel semiAnalytical Regressor using disenTangled Skeletal representations for human mesh recovery from videos, called ARTS. Specifically, a skeleton estimation and disentanglement module is proposed to estimate the 3D skeletons from a video and decouple them into disentangled skeletal representations (i.e., joint position, bone length, and human motion). Then, to fully utilize these representations, we introduce a semi-analytical regressor to estimate the parameters of the human mesh model. The regressor consists of three modules: Temporal Inverse Kinematics (TIK), Bone-guided Shape Fitting (BSF), and Motion-Centric Refinement (MCR). TIK utilizes joint position to estimate initial pose parameters and BSF leverages bone length to regress bone-aligned shape parameters. Finally, MCR combines human motion representation with image features to refine the initial human model parameters. Extensive experiments demonstrate that our ARTS surpasses existing state-of-the-art video-based methods in both per-frame accuracy and temporal consistency on popular benchmarks: 3DPW, MPI-INF-3DHP, and Human3.6M. Code is available at https://github.com/TangTao-PKU/ARTS.
Hong Liu 0008, Yingxuan You, Ti Wang, Wenhao Li 0002
ACM Multimedia2
2024 APP: Adaptive Pose Pooling for 3D Human Pose Estimation from Videos
Jinyan Zhang, Mengyuan Liu 0004, Hong Liu 0008, Wenhao Li 0002
ACM Multimedia3
2024 A gated cross-domain collaborative network for underwater object detection
Linhui Dai, Hong Liu 0008, Pinhao Song, Mengyuan Liu 0001
Pattern Recognit.2
2024 Improving self-supervised action recognition from extremely augmented skeleton sequences
Tianyu Guo 0001, Mengyuan Liu 0001, Hong Liu 0008, Wenhao Li 0002
Pattern Recognit.3
2024 Augmented skeleton sequences with hypergraph network for self-supervised group activity recognition
Hong Liu 0008, Peini Guo, Ti Wang, Jingwen Guo, Ruijia Fan
Pattern Recognit.3
2024 Neural ordinary differential equation for irregular human motion prediction
Hong Liu 0008, Pinhao Song, Wenhao Li 0002
Pattern Recognit. Lett.2
2024 MVSSC: Meta-reinforcement learning based visual indoor navigation using multi-view semantic spatial context
abstract
In Visual Indoor Navigation (VIN), Deep Reinforcement Learning (DRL) is commonly used by agents to achieve end-to-end mapping from vision to action when navigating toward a target based on observation. However, current DRL-based work suffers from two challenges: partial observability resulting from using solely a single first-person view and poor generalization in the case of unknown scenes and unknown objects. In light of these issues, this paper introduces the integration of multi-view as an expansion of observability and meta-learning as a primary generalization technique into the DRL framework and presents the meta-reinforcement learning method that leverages Multi-View Semantic Spatial Context (MVSSC). Specifically, aiming to explore the informative multi-view context for better efficiency in searching for and navigating to the target, we model the objects’ relationship from two aspects, multi-view semantic context (MVSEC) and multi-view spatial context (MVSPC). MVSEC enables agents to encode prior semantic relationships adaptively via a multi-view modulated graph. Meanwhile, MVSPC enhances the spatial representation of target-related objects’ correlation through similarity grids of multi-view. After adaptively fusing the multi-view and context information under the meta-reinforcement learning framework, our method can encourage efficient target search and robust navigation with stronger generalization performance to unknown scenes and unknown objects. Extensive experimental results on the AI2-THOR simulator demonstrate that our method outperforms the current state-of-the-art approaches.
Wanruo Zhang, Hong Liu 0008, Jianbing Wu, Yidi Li 0001
Pattern Recognit. Lett.2
2024 MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human Motions
abstract
In this paper, we address the unexplored question of temporal sentence localization in human motions (TSLM), aiming to locate a target moment from a 3D human motion that semantically corresponds to a text query. Considering that 3D human motions are captured using specialized motion capture devices, motions with only a few joints lack complex scene information like objects and lighting. Due to this character, motion data has low contextual richness and semantic ambiguity between frames, which limits the accuracy of predictions made by current video localization frameworks extended to TSLM to only a rough level. To refine this, we devise two novel label-prior-assisted training schemes: one embed prior knowledge of foreground and background to highlight the localization chances of target moments, and the other forces the originally rough predictions to overlap with the more accurate predictions obtained from the flipped start/end prior label sequences during recovery training. We show that injecting label-prior knowledge into the model is crucial for improving performance at high IoU. In our constructed TSLM benchmark, our model termedMLPachieves a recall of 44.13 at [email protected] on the BABEL dataset and 71.17 on HumanML3D (Restore), outperforming prior works. Finally, we showcase the potential of our approach in corpus-level moment retrieval. Our source code is openly accessible athttps://github.com/eanson023/mlp.
Mengyuan Liu 0001, Yong Wang 0053, Yang Liu 0264, Hong Liu 0008
IEEE Trans. Circuits Syst. Video Technol.5
2024 Cross-Model Cross-Stream Learning for Self-Supervised Human Action Recognition
abstract
Considering the instance-level discriminative ability, contrastive learning methods, including MoCo and SimCLR, have been adapted from the original image representation learning task to solve the self-supervised skeleton-based action recognition task. These methods usually use multiple data streams (i.e., joint, motion, and bone) for ensemble learning, meanwhile, how to construct a discriminative feature space within a single stream and effectively aggregate the information from multiple streams remains an open problem. To this end, this article first applies a new contrastive learning method called bootstrap your own latent (BYOL) to learn from skeleton data, and then formulate SkeletonBYOL as a simple yet effective baseline for self-supervised skeleton-based action recognition. Inspired by SkeletonBYOL, this article further presents a cross-model and cross-stream (CMCS) framework. This framework combines cross-model adversarial learning (CMAL) and cross-stream collaborative learning (CSCL). Specifically, CMAL learns single-stream representation by cross-model adversarial loss to obtain more discriminative features. To aggregate and interact with multistream information, CSCL is designed by generating similarity pseudolabel of ensemble learning as supervision and guiding feature generation for individual streams. Extensive experiments on three datasets verify the complementary properties between CMAL and CSCL and also verify that the proposed method can achieve better results than state-of-the-art methods using various evaluation protocols.
Mengyuan Liu 0001, Hong Liu 0008, Tianyu Guo 0001
IEEE Trans. Hum. Mach. Syst.2
2024 Feature Completion Transformer for Occluded Person Re-Identification
abstract
Occluded person re-identification is a challenging problem due to the destruction of occluders in different camera views. Most existing paradigms focus on visible human body parts through some external models to reduce noise interference. However, the feature misalignment problem caused by discarded occlusions negatively affects the performance of the network. Different from most previous works that discard the occluded regions, we present Feature Completion Transformer (FCFormer) that reduces noise interference and complements missing features in occluded parts. Specifically, Occlusion Instance Augmentation is proposed to simulate real and diverse occlusion situations on the holistic image, which enlarges the occlusion samples in the training set and forms aligned occluded-holistic pairs. To reduce the interference of noise, a two-stream architecture is proposed to learn pairwise discriminative features from aligned image pairs, while obtaining self-aligned occluded-holistic feature level sample-label pairs without additional auxiliary models. To complement the features of occluded regions, a Feature Completion Decoder is designed to aggregate possible information from self-generated occluded features in a self-supervised manner. Further, in order to correlate the completion features with identity information, Feature Completion Consistency loss is introduced to enforce the distribution of the generated completion features to be consistent with the real holistic feature distribution. In addition, we propose the Cross Hard Triplet loss to further bridge the gap between completion features and extracting features under the same ID. Extensive experiments over five challenging datasets demonstrate that the proposed FCFormer achieves superior performance and outperforms the state-of-theart methods by significant margins on Occluded-Duke dataset.
Mengyuan Liu 0001, Hong Liu 0008, Wenhao Li 0002, Miaoju Ban, Tianyu Guo 0001, Yidi Li 0001
IEEE Trans. Multim.3
2024 Style-Agnostic Representation Learning for Visible-Infrared Person Re-Identification
abstract
One main challenge of visible-infrared person re-identification (VI Re-ID) lies in the large style discrepancy between the heterogeneous data. We present a STyle-Agnostic Representation learning (STAR) framework that bridges the modality gaps at both data and feature levels in a progressive manner. At the data level, we present Cross Modality Blending (CMB), a powerful and parameter-free data augmentation scheme that smoothly synthesizes intermediate modalities by conducting identity-preserving patch exchange and smooth cross-modality blending. At the feature level, we explore the inter-modality feature alignment problem from a new perspective of the style-related feature statistics. Specifically, we design a plug-and-play Adaptive Style Normalization (ASN) module to discard the intrinsic style distractors without losing discriminative content via dual-level adaptive distribution normalization and discriminability compensation. Moreover, considering that an appropriate modality intermediary can convey relevant information on the inter-modality distribution shift, we propose Reciprocal Modality Bridging Learning (RMBL) to better steer the modality bridging process. Two lightweight modality transformation modules are designed in RMBL to model an appropriate intermediate space by manipulating high-order statistics under our shortest distance constraint. Meanwhile, intermediary-guided distribution alignment is reciprocally conducted to align heterogeneous features to the modality intermediary. Experiments on VI Re-ID benchmarks demonstrate the superiority and flexibility of STAR over state-of-the-art methods.
Jianbing Wu, Hong Liu 0008, Wei Shi 0009, Mengyuan Liu 0001, Wenhao Li 0002
IEEE Trans. Multim.2
2024 BCAN: Bidirectional Correct Attention Network for Cross-Modal Retrieval
abstract
As a fundamental topic in bridging the gap between vision and language, cross-modal retrieval purposes to obtain the correspondences' relationship between fragments, i.e., subregions in images and words in texts. Compared with earlier methods that focus on learning the visual semantic embedding from images and sentences to the shared embedding space, the existing methods tend to learn the correspondences between words and regions via cross-modal attention. However, such attention-based approaches invariably result in semantic misalignment between subfragments for two reasons: 1) without modeling the relationship between subfragments and the semantics of the entire images or sentences, it will be hard for such approaches to distinguish images or sentences with multiple same semantic fragments and 2) such approaches focus attention evenly on all subfragments, including nonvisual words and a lot of redundant regions, which also will face the problem of semantic misalignment. To solve these problems, this article proposes a bidirectional correct attention network (BCAN), which introduces a novel concept of the relevance between subfragments and the semantics of the entire images or sentences and designs a novel correct attention mechanism by modeling the local and global similarity between images and sentences to correct the attention weights focused on the wrong fragments. Specifically, we introduce a concept about the semantic relationship between subfragments and entire images or sentences and use this concept to solve the semantic misalignment from two aspects. In our correct attention mechanism, we design two independent units to correct the weight of attention focused on the wrong fragments. Global correct unit (GCU) with modeling the global similarity between images and sentences into the attention mechanism to solve the semantic misalignment problem caused by focusing attention on relevant subfragments in irrelevant pairs (RI) and the local correct unit (LCU) consider the difference in the attention weights between fragments among two steps to solve the semantic misalignment problem caused by focusing attention on irrelevant subfragments in relevant pairs (IR). Extensive experiments on large-scale MS-COCO and Flickr30K show that our proposed method outperforms all the attention-based methods and is competitive to the state-of-the-art. Our code and pretrained model are publicly available at: https://github.com/liuyyy111/BCAN.
Yang Liu 0264, Hong Liu 0008, Huaqiu Wang, Fanyang Meng, Mengyuan Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 HTNet: Human Topology aware network for 3d Human pose estimation
abstract
3D human pose estimation errors would propagate along the human body topology and accumulate at the end joints of limbs. Inspired by the backtracking mechanism in automatic control systems, we design an Intra-Part Constraint module that utilizes the parent nodes as the reference to build topological constraints for end joints at the part level. Further considering the hierarchy of the human topology, joint-level and body-level dependencies are captured via graph convolutional networks and self-attentions, respectively. Based on these designs, we propose a novel Human Topology aware Network (HTNet), which adopts a channel-split progressive strategy to sequentially learn the structural priors of the human topology from multiple semantic levels: joint, part, and body. Extensive experiments show that the proposed method improves the estimation accuracy by 18.7% on the end joints of limbs and achieves state-of-the-art results on Human3.6M and MPI-INF-3DHP datasets. Code is available at https://github.com/vefalun/HTNet.
Jialun Cai, Hong Liu 0008, Runwei Ding, Wenhao Li 0002, Jianbing Wu, Miaoju Ban
ICASSP2
2023 Body Prior Guided Graph Convolutional Neural Network for Skeleton-Based Action Recognition
abstract
Graph Convolutional Network (GCN) has achieved high success in the skeleton-based human action recognition task by modeling the human skeleton as a graph. However, it remains a problem for GCN-based methods to learn distinctive action features from a limited number of training samples. Via taking full advantage of the body prior knowledge, this paper presents a Body Prior Guided Graph Convolutional Network (BPG-GCN) to jointly meet the demand for large-scale training data and effective model architecture. Unlike standard GCN-based methods, our BPG-GCN additionally involves both Body Prior Guided Drop (BPGD) and Body Prior Guided Attention (BPGA) modules. Specifically, the BPGD module generates diverse augmented skeleton sequences by selectively dropping spatial-temporal skeleton joints. Moreover, the BPGA module combines body structure and attention mechanism to learn distinctive action features for specific body parts. Extensive experiments on NTU-60 and NW-UCLA datasets consistently verify the effectiveness of our proposed BPG-GCN by outperforming state-of-the-art GCN-based methods. Our code is publicly available at https://github.com/519542630/BPG-GCN.
Qianshuo Hu, Hong Liu 0008, Huaqiu Wang
ICASSP2
2023 Boosting Person Re-Identification with Viewpoint Contrastive Learning and Adversarial Training
abstract
Person re-identification (ReID) aims at retrieving a person of interest across multiple cameras. Despite significant progress in person ReID, viewpoint variation remains an obstacle to extracting discriminative features for retrieval. To address this problem, we propose a Viewpoint-Robust Network (VRN) based on contrastive learning and adversarial training to boost person ReID. Specifically, a View-point Confusion (VC) module is proposed to conceal viewpoint information to extract viewpoint-agnostic features. We employ viewpoint contrastive learning to discriminate viewpoints, and then conversely ignore the viewpoint information by adversarial training. Besides, an ID Prototype (IDP) module further enhances the network by introducing a confidence-weighted IDP as a viewpoint-robust ID representation and conducting contrastive metric learning with an IDP triplet loss. Extensive experiments demonstrate the proposed method achieves state-of-the-art performance on widely used datasets Market1501 and MSMT17. Visualization of retrieval results illustrates the effectiveness and robustness of the proposed method.
Xingyue Shi, Hong Liu 0008, Wei Shi 0009, Zihui Zhou, Yidi Li 0001
ICASSP2
2023 Interweaved Graph and Attention Network for 3D Human Pose Estimation
abstract
Despite substantial progress in 3D human pose estimation from a single-view image, prior works rarely explore global and local correlations, leading to insufficient learning of human skeleton representations. To address this issue, we propose a novel Interweaved Graph and Attention Network (IGANet) that allows bidirectional communications between graph convolutional networks (GCNs) and attentions. Specifically, we introduce an IGA module, where attentions are provided with local information from GCNs and GCNs are injected with global information from attentions. Additionally, we design a simple yet effective U-shaped multi-layer perceptron (uMLP), which can capture multi-granularity information for body joints. Extensive experiments on two popular benchmark datasets (i.e. Human3.6M and MPI-INF-3DHP) are conducted to evaluate our proposed method. The results show that IGANet achieves state-of-the-art performance on both datasets. Code is available at https://github.com/xiu-cs/IGANet.
Ti Wang, Hong Liu 0008, Runwei Ding, Wenhao Li 0002, Yingxuan You, Xia Li 0005
ICASSP2
2023 Gator: Graph-Aware Transformer with Motion-Disentangled Regression for Human Mesh Recovery from a 2D Pose
abstract
3D human mesh recovery from a 2D pose plays an important role in various applications. However, it is hard for existing methods to simultaneously capture the multiple relations during the evolution from skeleton to mesh, including joint-joint, joint-vertex and vertex-vertex relations, which often leads to implausible results. To address this issue, we propose a novel solution, called GATOR, that contains an encoder of Graph-Aware Transformer (GAT) and a decoder with Motion-Disentangled Regression (MDR) to explore these multiple relations. Specifically, GAT combines a GCN and a graph-aware self-attention in parallel to capture physical and hidden joint-joint relations. Furthermore, MDR models joint-vertex and vertex-vertex interactions to explore joint and vertex relations. Based on the clustering characteristics of vertex offset fields, MDR regresses the vertices by composing the predicted base motions. Extensive experiments show that GATOR achieves state-of-the-art performance on two challenging benchmarks. Code is available at https://github.com/kasvii/GATOR.
Yingxuan You, Hong Liu 0008, Xia Li 0005, Wenhao Li 0002, Ti Wang, Runwei Ding
ICASSP2
2023 FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge Distillation
abstract
Existing skeleton-based action recognition methods typically follow a centralized learning paradigm, which can pose privacy concerns when exposing human-related videos. Federated Learning (FL) has attracted much attention due to its outstanding advantages in privacy-preserving. However, directly applying FL approaches to skeleton videos suffers from unstable training. In this paper, we investigate and discover that the heterogeneous human topology graph structure is the crucial factor hindering training stability. To address this limitation, we pioneer a novel Federated Skeleton-based Action Recognition (FSAR) paradigm, which enables the construction of a globally generalized model without accessing local sensitive data. Specifically, we introduce an Adaptive Topology Structure (ATS), separating generalization and personalization by learning a domain-invariant topology shared across clients and a domain-specific topology decoupled from global model aggregation. Furthermore, we explore Multi-grain Knowledge Distillation (MKD) to mitigate the discrepancy between clients and server caused by distinct updating patterns through aligning shallow block-wise motion features. Extensive experiments on multiple datasets demonstrate that FSAR outperforms state-of-the-art FL-based methods while inherently protecting privacy.
Jingwen Guo, Hong Liu 0008, Shitong Sun, Tianyu Guo 0001, Min Zhang 0005, Chenyang Si
ICCV2
2023 Learning Concordant Attention via Target-aware Alignment for Visible-Infrared Person Re-identification
abstract
Owing to the large distribution gap between the heterogeneous data in Visible-Infrared Person Re-identification (VI Re-ID), we point out that existing paradigms often suffer from the inter-modal semantic misalignment issue and thus fail to align and compare local details properly. In this paper, we present Concordant Attention Learning (CAL), a novel framework that learns semantic-aligned representations for VI Re-ID. Specifically, we design the Target-aware Concordant Alignment paradigm, which allows target-aware attention adaptation when aligning heterogeneous samples (i.e., adaptive attention adjustment according to the target image being aligned). This is achieved by exploiting the discriminative clues from the modality counterpart and designing effective modality-agnostic correspondence searching strategies. To ensure semantic concordance during the cross-modal retrieval stage, we further propose MatchDistill, which matches the attention patterns across modalities and learns their underlying semantic correlations by bipartite-graph-based similarity modeling and cross-modal knowledge exchange. Extensive experiments on VI Re-ID benchmark datasets demonstrate the effectiveness and superiority of the proposed CAL.
Jianbing Wu, Hong Liu 0008, Yuxin Su 0004, Wei Shi 0009, Hao Tang 0005
ICCV2
2023 Co-Evolution of Pose and Mesh for 3D Human Body Estimation from Video
abstract
Despite significant progress in single image-based 3D human mesh recovery, accurately and smoothly recovering 3D human motion from a video remains challenging. Existing video-based methods generally recover human mesh by estimating the complex pose and shape parameters from coupled image features, whose high complexity and low representation ability often result in inconsistent pose motion and limited shape patterns. To alleviate this issue, we introduce 3D pose as the intermediary and propose a Pose and Mesh Co-Evolution network (PMCE) that decouples this task into two parts: 1) video-based 3D human pose estimation and 2) mesh vertices regression from the estimated 3D pose and temporal image feature. Specifically, we propose a two-stream encoder that estimates mid-frame 3D pose and extracts a temporal image feature from the input image sequence. In addition, we design a co-evolution decoder that performs pose and mesh interactions with the image-guided Adaptive Layer Normalization (AdaLN) to make pose and mesh fit the human body shape. Extensive experiments demonstrate that the proposed PMCE outperforms previous state-of-the-art methods in terms of both per-frame accuracy and temporal consistency on three benchmark datasets: 3DPW, Human3.6M, and MPI-INF-3DHP. Our code is available at https://github.com/kasvii/PMCE.
Yingxuan You, Hong Liu 0008, Ti Wang, Wenhao Li 0002, Runwei Ding, Xia Li 0005
ICCV2
2023 Self-Supervised 3D Skeleton Representation Learning with Active Sampling and Adaptive Relabeling for Action Recognition
abstract
Self-supervised 3D skeleton representation learning has recently shown great potential for action recognition via contrastive learning. However, existing methods suffer from limited learning efficiency and the unreliability of representations, which is not conducive to action recognition. To this end, we propose an Active Sampling and Adaptive Relabeling (ASAR) contrastive learning method to achieve efficient and reliable learning of 3D skeleton representations. Specifically, the active sampling strategy is used to build a dictionary with informative samples for efficient representation learning. Additionally, the adaptive relabeling strategy is proposed to automatically modify the confidence scores of the extra positive samples and alleviate the unreliability of representations. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets demonstrate the superiority of our approach.
Hong Liu 0008, Tianyu Guo 0001, Jingwen Guo, Ti Wang, Yidi Li 0001
ICIP2
2023 Semantic-aware Consistency Network for Cloth-changing Person Re-Identification
abstract
Cloth-changing Person Re-Identification (CC-ReID) is a challenging task that aims to retrieve the target person across multiple surveillance cameras when clothing changes might happen. Despite recent progress in CC-ReID, existing approaches are still hindered by the interference of clothing variations since they lack effective constraints to keep the model consistently focused on clothing-irrelevant regions. To address this issue, we present a Semantic-aware Consistency Network (SCNet) to learn identity-related semantic features by proposing effective consistency constraints. Specifically, we generate the black-clothing image by erasing pixels in the clothing area, which explicitly mitigates the interference from clothing variations. In addition, to fully exploit the fine-grained identity information, a head-enhanced attention module is introduced, which learns soft attention maps by utilizing the proposed part-based matching loss to highlight head information. We further design a semantic consistency loss to facilitate the learning of high-level identity-related semantic features, forcing the model to focus on semantically consistent cloth-irrelevant regions. By using the consistency constraint, our model does not require any extra auxiliary segmentation module to generate the black-clothing image or locate the head region during the inference stage. Extensive experiments on four cloth-changing person Re-ID datasets (LTCC, PRCC, Vc-Clothes, and DeepChange) demonstrate that our proposed SCNet makes significant improvements over prior state-of-the-art approaches. Our code is available at: https://github.com/Gpn-star/SCNet.
Peini Guo, Hong Liu 0008, Jianbing Wu
ACM Multimedia2
2023 Cross-Modal Retrieval for Motion and Text via DropTriple Loss
abstract
Cross-modal retrieval of image-text and video-text is a prominent research area in computer vision and natural language processing. However, there has been insufficient attention given to cross-modal retrieval between human motion and text, despite its wide-ranging applicability. To address this gap, we utilize a concise yet effective dual-unimodal transformer encoder for tackling this task. Recognizing that overlapping atomic actions in different human motion sequences can lead to semantic conflicts between samples, we explore a novel triplet loss function called DropTriple Loss. This loss function discards false negative samples from the negative sample set and focuses on mining remaining genuinely hard negative samples for triplet training, thereby reducing violations they cause. We evaluate our model and approach on the HumanML3D and KIT Motion-Language datasets. On the latest HumanML3D dataset, we achieve a recall of 62.9% for motion retrieval and 71.5% for text retrieval (both based on R@10). The source code for our approach is publicly available at https://github.com/eanson023/rehamot.
Yang Liu 0264, Haoqiang Wang, Mengyuan Liu 0001, Hong Liu 0008
MMAsia6
2023 Achieving domain generalization for underwater object detection by domain mixup and contrastive learning
Pinhao Song, Hong Liu 0008, Linhui Dai, Xiaochuan Zhang, Runwei Ding, Shengquan Li 0001
Neurocomputing3
2023 Multi-hypothesis representation learning for transformer-based 3D human pose estimation
abstract
Despite significant progress, estimating 3D human poses from monocular videos remains a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasible solutions ( i.e. , hypotheses) exist. To relieve this limitation, we propose a Multi-Hypothesis Transformer that learns spatio-temporal representations of multiple plausible pose hypotheses. In order to effectively model multi-hypothesis dependencies and build strong relationships across hypothesis features, we introduce a one-to-many-to-one three-stage framework: (i) Generate multiple initial hypothesis representations; (ii) Model self-hypothesis communication, merge multiple hypotheses into a single converged representation and then partition it into several diverged hypotheses; (iii) Learn cross-hypothesis communication and aggregate the multi-hypothesis features to synthesize the final 3D pose. Through the above processes, the final representation is enhanced and the synthesized pose is much more accurate. Extensive experiments show that the proposed method achieves state-of-the-art results on two challenging datasets: Human3.6M and MPI-INF-3DHP. The code and models are available at https://github.com/Vegetebird/MHFormer .
Wenhao Li 0002, Hong Liu 0008, Hao Tang 0005, Pichao Wang
Pattern Recognit.2
2023 AO2-DETR: Arbitrary-Oriented Object Detection Transformer
abstract
Arbitrary-oriented object detection (AOOD) is a challenging task to detect objects in the wild with arbitrary orientations and cluttered arrangements. Existing approaches are mainly based on anchor-based boxes or dense points, which rely on complicated hand-designed processing steps and inductive bias, such as anchor generation, transformation, and non-maximum suppression reasoning. Recently, the emerging transformer-based approaches view object detection as a direct set prediction problem that effectively removes the need for hand-designed components and inductive biases. In this paper, we propose an Arbitrary-Oriented Object DEtection TRansformer framework, termed AO2-DETR, which comprises three dedicated components. More precisely, an oriented proposal generation mechanism is proposed to explicitly generate oriented proposals, which provides better positional priors for pooling features to modulate the cross-attention in the transformer decoder. An adaptive oriented proposal refinement module is introduced to extract rotation-invariant region features and eliminate the misalignment between region features and objects. And a rotation-aware set matching loss is used to ensure the one-to-one matching process for direct set prediction without duplicate predictions. Our method considerably simplifies the overall pipeline and presents a new AOOD paradigm. Comprehensive experiments on several challenging datasets show that our method achieves superior performance on the AOOD task.
Linhui Dai, Hong Liu 0008, Hao Tang 0005, Pinhao Song
IEEE Trans. Circuits Syst. Video Technol.2
2023 Multi-Dimensional Attention With Similarity Constraint for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) is a challenging task in understanding untrimmed videos, in which no frame-wise annotation is provided during training, only the video-level category label is available. Current methods mainly adopt temporal attention branches to conduct foreground-background separation with RGB and optical flow features simply concatenated, regardless of the discriminative spacial features and the complementarity between different modalities. In this work, we propose a Multi-Dimensional Attention (MDA) method to explore attention mechanism across three dimensions in weakly supervised action localization,i.e., 1) temporal attention that focuses on segments containing action instances, 2) channel attention that discovers the most relevant cues for action description, and 3) modal attention that fuses RGB and flow information adaptively based on feature magnitudes during background modeling. In addition, we introduce a similarity constraint loss to refine the action segment representation in feature space, which helps the network to detect less discriminative frames of an action to capture the full action boundaries. The proposed MDA with similarity constraints can be easily applied to existing action detection frameworks with few parameters. Extensive experiments on THUMOS’14 and ActivityNet v1.2 datasets show that the proposed method outperforms the current state-of-the-art WTAL approaches, and achieves comparable results with some advanced fully-supervised methods.
Zhengyan Chen, Hong Liu 0008, Xin Liao 0001
IEEE Trans. Multim.2
2023 Weakly-Supervised 3D Human Pose Estimation With Cross-View U-Shaped Graph Convolutional Network
abstract
Although monocular 3D human pose estimation methods have made significant progress, it is far from being solved due to the inherent depth ambiguity. Instead, exploiting multi-view information is a practical way to achieve absolute 3D human pose estimation. In this paper, we propose a simple yet effective pipeline for weakly-supervised cross-view 3D human pose estimation. By only using two camera views, our method can achieve state-of-the-art performance in a weakly-supervised manner, requiring no 3D ground truth but only 2D annotations. Specifically, our method contains two steps: triangulation and refinement. First, given the 2D keypoints that can be obtained through any classic 2D detection methods, triangulation is performed across two views to lift the 2D keypoints into coarse 3D poses. Then, a novel cross-view U-shaped graph convolutional network (CV-UGCN), which can explore the spatial configurations and cross-view correlations, is designed to refine the coarse 3D poses. In particular, the refinement progress is achieved through weakly-supervised learning, in which geometric and structure-aware consistency checks are performed. We evaluate our method on the standard benchmark dataset, Human3.6M. The Mean Per Joint Position Error on the benchmark dataset is 27.4 mm, which outperforms existing state-of-the-art methods remarkably (27.4 mm vs 30.2 mm).
Guoliang Hua, Hong Liu 0008, Wenhao Li 0002, Runwei Ding, Xin Xu 0001
IEEE Trans. Multim.2
2023 Exploiting Temporal Contexts With Strided Transformer for 3D Human Pose Estimation
abstract
Despite the great progress in 3D human pose estimation from videos, it is still an open problem to take full advantage of a redundant 2D pose sequence to learn representative representations for generating one 3D pose. To this end, we propose an improved Transformer-based architecture, called Strided Transformer, which simply and effectively lifts a long sequence of 2D joint locations to a single 3D pose. Specifically, a Vanilla Transformer Encoder (VTE) is adopted to model long-range dependencies of 2D pose sequences. To reduce the redundancy of the sequence, fully-connected layers in the feed-forward network of VTE are replaced with strided convolutions to progressively shrink the sequence length and aggregate information from local contexts. The modified VTE is termed as Strided Transformer Encoder (STE), which is built upon the outputs of VTE. STE not only effectively aggregates long-range information to a single-vector representation in a hierarchical global and local fashion, but also significantly reduces the computation cost. Furthermore, a full-to-single supervision scheme is designed at both full sequence and single target frame scales applied to the outputs of VTE and STE, respectively. This scheme imposes extra temporal smoothness constraints in conjunction with the single target frame supervision and hence helps produce smoother and more accurate 3D poses. The proposed Strided Transformer is evaluated on two challenging benchmark datasets, Human3.6 M and HumanEva-I, and achieves state-of-the-art results with fewer parameters. Code and models are available athttps://github.com/Vegetebird/StridedTransformer-Pose3D.
Wenhao Li 0002, Hong Liu 0008, Runwei Ding, Mengyuan Liu 0001, Pichao Wang, Wenming Yang
IEEE Trans. Multim.2
2023 AttentionGAN: Unpaired Image-to-Image Translation Using Attention-Guided Generative Adversarial Networks
abstract
State-of-the-art methods in the image-to-image translation are capable of learning a mapping from a source domain to a target domain with unpaired image data. Though the existing methods have achieved promising results, they still produce visual artifacts, being able to translate low-level information but not high-level semantics of input images. One possible reason is that generators do not have the ability to perceive the most discriminative parts between the source and target domains, thus making the generated images low quality. In this article, we propose a new Attention-Guided Generative Adversarial Networks (AttentionGAN) for the unpaired image-to-image translation task. AttentionGAN can identify the most discriminative foreground objects and minimize the change of the background. The attention-guided generators in AttentionGAN are able to produce attention masks, and then fuse the generation output with the attention masks to obtain high-quality target images. Accordingly, we also design a novel attention-guided discriminator which only considers attended regions. Extensive experiments are conducted on several generative tasks with eight public datasets, demonstrating that the proposed method is effective to generate sharper and more realistic images compared with existing competitive models. The code is available at https://github.com/Ha0Tang/AttentionGAN.
Hao Tang 0005, Hong Liu 0008, Dan Xu 0002, Philip Torr 0001, Nicu Sebe
IEEE Trans. Neural Networks Learn. Syst.2
2022 Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action Recognition
abstract
In recent years, self-supervised representation learning for skeleton-based action recognition has been developed with the advance of contrastive learning methods. The existing contrastive learning methods use normal augmentations to construct similar positive samples, which limits the ability to explore novel movement patterns. In this paper, to make better use of the movement patterns introduced by extreme augmentations, a Contrastive Learning framework utilizing Abundant Information Mining for self-supervised action Representation (AimCLR) is proposed. First, the extreme augmentations and the Energy-based Attention-guided Drop Module (EADM) are proposed to obtain diverse positive samples, which bring novel movement patterns to improve the universality of the learned representations. Second, since directly using extreme augmentations may not be able to boost the performance due to the drastic changes in original identity, the Dual Distributional Divergence Minimization Loss (D3M Loss) is proposed to minimize the distribution divergence in a more gentle way. Third, the Nearest Neighbors Mining (NNM) is proposed to further expand positive samples to make the abundant information mining process more reasonable. Exhaustive experiments on NTU RGB+D 60, PKU-MMD, NTU RGB+D 120 datasets have verified that our AimCLR can significantly perform favorably against state-of-the-art methods under a variety of evaluation protocols with observed higher quality action representations. Our code is available at https://github.com/Levigty/AimCLR.
Tianyu Guo 0001, Hong Liu 0008, Mengyuan Liu 0001, Runwei Ding
AAAI2
2022 Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker Tracking
abstract
Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of multi-modal signals remains a challenging issue. In this paper, we propose a novel Multi-modal Perception Tracker (MPT) for speaker tracking using both audio and visual modalities. Specifically, a novel acoustic map based on spatial-temporal Global Coherence Field (stGCF) is first constructed for heterogeneous signal fusion, which employs a camera model to map audio cues to the localization space consistent with the visual cues. Then a multi-modal perception attention network is introduced to derive the perception weights that measure the reliability and effectiveness of intermittent audio and video streams disturbed by noise. Moreover, a unique cross-modal self-supervised learning method is presented to model the confidence of audio and visual observations by leveraging the complementarity and consistency between different modalities. Experimental results show that the proposed MPT achieves 98.6% and 78.3% tracking accuracy on the standard and occluded datasets, respectively, which demonstrates its robustness under adverse conditions and outperforms the current state-of-the-art methods.
Yidi Li 0001, Hong Liu 0008, Hao Tang 0005
AAAI2
2022 Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on Transformer
abstract
Occluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem by aligning body parts according to graph matching, but these graph-based methods are not intuitive and complicated. Therefore, we propose a transformer-based Pose-guided Feature Disentangling (PFD) method by utilizing pose information to clearly disentangle semantic components (e.g. human body or joint parts) and selectively match non-occluded parts correspondingly. First, Vision Transformer (ViT) is used to extract the patch features with its strong capability. Second, to preliminarily disentangle the pose information from patch information, the matching and distributing mechanism is leveraged in Pose-guided Feature Aggregation (PFA) module. Third, a set of learnable semantic views are introduced in transformer decoder to implicitly enhance the disentangled body part features. However, those semantic views are not guaranteed to be related to the body without additional supervision. Therefore, Pose-View Matching (PVM) module is proposed to explicitly match visible body parts and automatically separate occlusion features. Fourth, to better prevent the interference of occlusions, we design a Pose-guided Push Loss to emphasize the features of visible body parts. Extensive experiments over five challenging datasets for two tasks (occluded and holistic Re-ID) demonstrate that our proposed PFD is superior promising, which performs favorably against state-of-the-art methods. Code is available at https://github.com/WangTaoAs/PFD_Net
Hong Liu 0008, Pinhao Song, Tianyu Guo 0001, Wei Shi 0009
AAAI2
2022 MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
abstract
Estimating 3D human poses from monocular videos is a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasible solutions (i.e., hypotheses) exist. To relieve this limitation, we propose a Multi-Hypothesis Transformer (MHFormer) that learns spatio-temporal representations of multiple plausible pose hypotheses. In order to effectively model multi-hypothesis dependencies and build strong relationships across hypothesis features, the task is decomposed into three stages: (i) Generate multiple initial hypothesis representations; (ii) Model self-hypothesis communication, merge multiple hypotheses into a single converged representation and then partition it into several diverged hypotheses; (iii) Learn cross-hypothesis communication and aggregate the multi-hypothesis features to synthesize the final 3D pose. Through the above processes, the final representation is enhanced and the synthesized pose is much more accurate. Extensive experiments show that MHFormer achieves state-of-the-art results on two challenging datasets: Human3.6M and MPI-INF-3DHP. Without bells and whistles, its performance surpasses the previous best result by a large margin of 3% on Human3.6M. Code and models are available at https://github.com/Vegetebird/MHFormer.
Wenhao Li 0002, Hong Liu 0008, Hao Tang 0005, Pichao Wang, Luc Van Gool
CVPR2
2022 Adaptive Weighted Network With Edge Enhancement Module For Monocular Self-Supervised Depth Estimation
abstract
Monocular self-supervised depth estimation can be easily applied in many areas since only a single camera is required. However, current methods do not predict well in depth borders. Besides, factors such as occlusion and texture sparsity can lead to the failure of the photometric consistency, affecting the prediction performance. To overcome these deficiencies, an adaptive weighted monocular self-supervised depth estimation framework that exploits enhanced edge information and texture sparsity based adaptive weights is proposed. In particular, a module named edge enhancement module (EEM) is designed to be embedded into the current depth prediction network to extract edge details for clearer depth prediction in depth borders. Moreover, a texture sparsity based adaptive weighted (TSAW) loss is introduced to as-sign different weights according to texture sparsity, enabling a more targeted construction of geometric constraints. Experimental results on the KITTI dataset demonstrate that the proposed network outperforms state-of-the-art methods.
Hong Liu 0008, Guoliang Hua, Weibo Huang, Runwei Ding
ICASSP1
2022 SRP-DNN: Learning Direct-Path Phase Difference for Multiple Moving Sound Source Localization
abstract
Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning techniques to learn competing and time-varying direct-path phase differences for localizing multiple moving sound sources. A causal convolutional recurrent neural network is designed to extract the direct-path phase difference sequence from signals of each microphone pair. To avoid the assignment ambiguity and the problem of uncertain output-dimension encountered when simultaneously predicting multiple targets, the learning target is designed in a weighted sum format, which encodes source activity in the weight and direct-path phase differences in the summed value. The learned direct-path phase differences for all microphone pairs can be directly used to construct the spatial spectrum according to the formulation of steered response power (SRP). This deep neural network (DNN) based SRP method is referred to as SRP-DNN. The locations of sources are estimated by iteratively detecting and removing the dominant source from the spatial spectrum, in which way the interaction between sources is reduced. Experimental results on both simulated and real-world data show the superiority of the proposed method in the presence of noise and reverberation.
Bing Yang 0004, Hong Liu 0008, Xiaofei Li 0001
ICASSP2
2022 Unsupervised Domain Adaptation Person Re-Identification by Camera-Aware Style Decoupling and Uncertainty Modeling
abstract
Unsupervised domain adaptation (UDA) person re-identification (re-ID) aims to transfer knowledge learned from labeled source domain to unlabeled target domain and has been successfully applied into a wide range of real-world scenarios. However, existing methods are mainly ineffective at handling domain shift as well as being sensitive to camera styles due to the unannotated target domain. In this paper, we pro-pose a Camera-style Separation and Uncertainty Estimation (CSUE) model to address the problem from two perspectives. To alleviate the negative effect of cross-camera variation, we introduce the Camera-aware Style Decoupling module to im-pose inter-and-intra camera constraints on the feature extracting stage. It can better mine and describe the latent camera invariant features. Moreover, to avoid the inherent defect of clustering, an Uncertainty Modeling module is constructed via estimating the certainty, which helps progressively refine the pseudo labels. Extensive experiments on widely used datasets demonstrate the state-of-the-art performance of our model under the UDA re-ID setting.
Jingwen Guo, Hong Liu 0008, Wei Shi 0009, Hao Tang 0005, Jianbing Wu
ICIP2
2022 Identity-Sensitive Knowledge Propagation for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID), which aims to match person identities under clothing changes, is a new rising research topic in recent years. However, typical biometrics-based CC-ReID methods often require cumber-some pose or body part estimators to learn cloth-irrelevant features from human biometric traits, which comes with high computational costs. Besides, the performance is significantly limited due to the resolution degradation of surveillance images. To address the above limitations, we propose an effective Identity-Sensitive Knowledge Propagation framework (DeSKPro) for CC-ReID. Specifically, a Cloth-irrelevant Spatial Attention module is introduced to eliminate the distraction of clothing appearance by acquiring knowledge from the human parsing module. To mitigate the resolution degradation issue and mine identity-sensitive cues from human faces, we propose to restore the missing facial details using prior facial knowledge, which is then propagated to a smaller network. After training, the extra computations for human parsing or face restoration are no longer required. Extensive experiments show that our framework outperforms state-of-the-art methods by a large margin. Our code is available at https://github.com/KimbingNg/DeskPro.
Jianbing Wu, Hong Liu 0008, Wei Shi 0009, Hao Tang 0005, Jingwen Guo
ICIP2
2022 A Cloth-Irrelevant Harmonious Attention Network for Cloth-Changing Person Re-identification
abstract
Cloth-changing person re-identification (CC-ReID) is a challenging task that aims at retrieving the target person across large spatial and temporal spans, with a high probability of changing clothes. To explicitly alleviate the impact of the person changing clothes on re-identification, this paper presents a cloth-irrelevant harmonious attention network (CIHANet) that learns cloth-irrelevant knowledge. Firstly, with the help of human parsing, the color information of human clothing is removed to generate black clothes images. Secondly, the raw person images are used to learn features with more color-based appearance knowledge, while the black clothes images are used to learn features with more cloth-irrelevant knowledge. Then, to fuse the knowledge of two distinct streams, we propose the harmonious attention module, including mutual learning attention and salience guided attention mechanisms. The mutual learning attention mechanism adaptively selects identity-relevant features across feature channels to make two streams interact with each other. The salience guided attention mechanism highlights the cloth-irrelevant areas by transferring the spatial knowledge from the black clothes stream to the raw images stream. Finally, quantitative and qualitative results on three CC-ReID datasets validate the superiority of our method on the CC-ReID task.
Zihui Zhou, Hong Liu 0008, Wei Shi 0009, Hao Tang 0005, Xingyue Shi
ICPR2
2022 Integrating Point and Line Features for Visual-Inertial Initialization
abstract
Accurate and robust initialization is crucial in visual-inertial system, which significantly affects the localization accuracy. Most of the existing feature-based initialization methods rely on point features to estimate initial parameters. However, the performance of these methods often decreases in real scene, as point features are unstable and may be discontinuously observed especially in low textured environments. By contrast, line features, providing richer geometrical information than points, are also very common in man-made buildings. Thereby, in this paper, we propose a novel visual-inertial initialization method integrating both point and line features. Specifically, a closed-form method of line features is presented for initialization, which is combined with point-based method to build an integrated linear system. Parameters including initial velocity, gravity, point depth and line's endpoints depth can be jointly solved out. Furthermore, to refine these parameters, a global optimization method is proposed, which consists of two novel nonlinear least squares problems for respective points and lines. Both gravity magnitude and gyroscope bias are considered in refinement. Extensive experimental results on both simulated and public datasets show that integrating point and line features in initialization stage can achieve higher accuracy and better robustness compared with pure point-based methods.
Hong Liu 0008, Junyin Qiu, Weibo Huang
ICRA1
2022 IRANet: Identity-relevance aware representation for cloth-changing person re-identification
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001
Image Vis. Comput.2
2022 Image-to-video person re-identification using three-dimensional semantic appearance alignment and cross-modal interactive learning
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001
Pattern Recognit.2
2022 Regularizing Visual Semantic Embedding With Contrastive Learning for Image-Text Matching
abstract
Learning visual semantic embedding for image-text matching has achieved high success by using triplet loss to pull positive image-text pairs which share similar semantic meaning and to push negative image-text pairs which share different semantic meaning. Without modeling constraints from image-image or text-text pairs, the generated visual semantic embedding inevitably faces the problem of semantic misalignments among similar images or among similar texts. To solve this problem, we present a contrastive visual semantic embedding framework, named as ConVSE, which achieves intra-modal semantic alignment by contrastive learning from augmented image-image (or text-text) pairs and achieves inter-modal semantic alignment by applying hardest-negative-enhanced triplet loss on image-text pairs. To the best of our knowledge, we are the first to find that contrastive learning benefits visual semantic embedding. Extensive experiments on large scale MSCOCO and Flickr30K datasets verify the effectiveness of our proposed ConVSE by outperforming visual semantic embedding-based methods and achieving new state-of-the-arts. Our code and pretrained model are publicly available at: \url{https://github.com/liuyyy111/ConVSE}.
Yang Liu 0264, Hong Liu 0008, Huaqiu Wang, Mengyuan Liu 0001
IEEE Signal Process. Lett.2
2021 Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
Graph convolutional networks have been widely used for skeleton-based action recognition due to their excellent modeling ability of non-Euclidean data. As the graph convolution is a local operation, it can only utilize the short-range joint dependencies and short-term trajectory but fails to directly model the distant joints relations and long-range temporal information that are vital to distinguishing various actions. To solve this problem, we present a multi-scale spatial graph convolution (MS-GC) module and a multi-scale temporal graph convolution (MT-GC) module to enrich the receptive field of the model in spatial and temporal dimensions. Concretely, the MS-GC and MT-GC modules decompose the corresponding local graph convolution into a set of sub-graph convolution, forming a hierarchical residual architecture. Without introducing additional parameters, the features will be processed with a series of sub-graph convolutions, and each node could complete multiple spatial and temporal aggregations with its neighborhoods. The final equivalent receptive field is accordingly enlarged, which is capable of capturing both short- and long-range dependencies in spatial and temporal domains. By coupling these two modules as a basic block, we further propose a multi-scale spatial temporal graph convolutional network (MST-GCN), which stacks multiple blocks to learn effective motion representations for action recognition. The proposed MST-GCN achieves remarkable performance on three challenging benchmark datasets, NTU RGB+D, NTU-120 RGB+D and Kinetics-Skeleton, for skeleton-based action recognition.
Sicheng Li 0003, Bing Yang 0004, Qinghan Li, Hong Liu 0008
AAAI5
2021 Supervised Direct-Path Relative Transfer Function Learning for Binaural Sound Source Localization
abstract
Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two channels. Though DP-RTF fully encodes the sound directional cues and serves as a reliable localization feature, it is often erroneously estimated in the presence of noise and reverberation. This paper proposes a supervised DP-RTF learning method with deep neural networks for robust binaural sound source localization. To exploit the complementarity of single-channel spectrogram and dual-channel difference information, we first recover the direct-path magnitude spectrogram from the contaminated one using a monaural enhancement network, and then predict the DP-RTF from the dual-channel (enhanced-) intensity and phase cues using a binaural enhancement network. In addition, a weighted-matching softmax training loss is designed to promote the predicted DP-RTFs to be concentrated for the same direction and separated for different directions. Finally, the direction of arrival (DOA) of source is estimated by matching the predicted DP-RTF with the ground truths of candidate directions. Experimental results show the superiority of our method for DOA estimation in the environments with various levels of noise and reverberation.
Bing Yang 0004, Xiaofei Li 0001, Hong Liu 0008
ICASSP3
2021 Attend, Correct And Focus: A Bidirectional Correct Attention Network For Image-Text Matching
abstract
Image-text matching task aims to learn the fine-grained correspondences between images and sentences. Existing methods use attention mechanism to learn the correspondences by attending to all fragments without considering the relationship between fragments and global semantics, which inevitably lead to semantic misalignment among irrelevant fragments. To this end, we propose a Bidirectional Correct Attention Network (BCAN), which leverages global similarities and local similarities to reassign the attention weight, to avoid such semantic misalignment. Specifically, we introduce a global correct unit to correct the attention focused on relevant fragments in irrelevant semantics. A local correct unit is used to correct the attention focused on irrelevant fragments in relevant semantics. Experiments on Flickr30K and MSCOCO datasets verify the effectiveness of our proposed BCAN by outperforming both previous attention-based methods and state-of-the-art methods. Code can be found at: https://github.com/liuyyy111/BCAN.
Yang Liu 0264, Huaqiu Wang, Fanyang Meng, Mengyuan Liu 0001, Hong Liu 0008
ICIP5
2021 Modality-aware Style Adaptation for RGB-Infrared Person Re-Identification
abstract
RGB-infrared (IR) person re-identification is a challenging task due to the large modality gap between RGB and IR images. Many existing methods bridge the modality gap by style conversion, requiring high-similarity images exchanged by complex CNN structures, like GAN. In this paper, we propose a highly compact modality-aware style adaptation (MSA) framework, which aims to explore more potential relations between RGB and IR modalities by introducing new related modalities. Therefore, the attention is shifted from bridging to filling the modality gap with no requirement on high-quality generated images. To this end, we firstly propose a concise feature-free image generation structure to adapt the original modalities to two new styles that are compatible with both inputs by patch-based pixel redistribution. Secondly, we devise two image style quantification metrics to discriminate styles in image space using luminance and contrast. Thirdly, we design two image-level losses based on the quantified results to guide the style adaptation during an end-to-end four-modality collaborative learning process. Experimental results on two datasets SYSU-MM01 and RegDB show that MSA achieves significant improvements with little extra computation cost and outperforms the state-of-the-art methods.
Ziling Miao, Hong Liu 0008, Wei Shi 0009, Wanlu Xu, Hanrong Ye
IJCAI2
2021 Adversarial Feature Disentanglement for Long-Term Person Re-identification
abstract
Most existing person re-identification methods are effective in short-term scenarios because of their appearance dependencies. However, these methods may fail in long-term scenarios where people might change their clothes. To this end, we propose an adversarial feature disentanglement network (AFD-Net) which contains intra-class reconstruction and inter-class adversary to disentangle the identity-related and identity-unrelated (clothing) features. For intra-class reconstruction, the person images with the same identity are represented and disentangled into identity and clothing features by two separate encoders, and further reconstructed into original images to reduce intra-class feature variations. For inter-class adversary, the disentangled features across different identities are exchanged and recombined to generate adversarial clothes-changing images for training, which makes the identity and clothing features more independent. Especially, to supervise these new generated clothes-changing images, a re-feeding strategy is designed to re-disentangle and reconstruct these new images for image-level self-supervision in the original image space and feature-level soft-supervision in the disentangled feature space. Moreover, we collect a challenging Market-Clothes dataset and a real-world PKU-Market-Reid dataset for evaluation. The results on one large-scale short-term dataset (Market-1501) and five long-term datasets (three public and two we proposed) confirm the superiority of our method against other state-of-the-art methods.
Wanlu Xu, Hong Liu 0008, Wei Shi 0009, Ziling Miao, Zhisheng Lu, Feihu Chen
IJCAI2
2021 PCLoss: Fashion Landmark Estimation with Position Constraint Loss
Meijia Song, Hong Liu 0008, Wei Shi 0009, Xia Li 0005
Pattern Recognit.2
2021 Learning Deep Direct-Path Relative Transfer Function for Binaural Sound Source Localization
abstract
Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two microphone channels. Though DP-RTF fully encodes the sound spatial cues and serves as a reliable localization feature, it is often erroneously estimated in the presence of noise and reverberation. This paper proposes to learn DP-RTF with deep neural networks for robust binaural sound source localization. A DP-RTF learning network is designed to regress the binaural sensor signals to a real-valued representation of DP-RTF. It consists of a branched convolutional neural network module to separately extract the inter-channel magnitude and phase patterns, and a convolutional recurrent neural network module for joint feature learning. To better explore the speech spectra to aid the DP-RTF estimation, a monaural speech enhancement network is used to recover the direct-path spectrograms from the noisy ones. The enhanced spectrograms are stacked onto the noisy spectrograms to act as the input of the DP-RTF learning network. We train one unique DP-RTF learning network using many different binaural arrays to enable the generalization of DP-RTF learning across arrays. This way avoids time-consuming training data collection and network retraining for a new array, which is very useful in practical application. Experimental results on both simulated and real-world data show the effectiveness of the proposed method for direction of arrival (DOA) estimation in the noisy and reverberant environment, and a good generalization ability to unseen binaural arrays.
Bing Yang 0004, Hong Liu 0008, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Bi-Directional Exponential Angular Triplet Loss for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared person re-identification (RGB-IR Re-ID) is a cross-modality matching problem, where the modality discrepancy is a big challenge. Most existing works use Euclidean metric based constraints to resolve the discrepancy between features of images from different modalities. However, these methods are incapable of learning angularly discriminative feature embedding because Euclidean distance cannot measure the included angle between embedding vectors effectively. As an angularly discriminative feature space is important for classifying the human images based on their embedding vectors, in this paper, we propose a novel ranking loss function, named Bi-directional Exponential Angular Triplet Loss, to help learn an angularly separable common feature space by explicitly constraining the included angles between embedding vectors. Moreover, to help stabilize and learn the magnitudes of embedding vectors, we adopt a common space batch normalization layer. The quantitative and qualitative experiments on the SYSU-MM01 and RegDB dataset support our analysis. On SYSU-MM01 dataset, the performance is improved from 7.40% / 11.46% to 38.57% / 38.61% for rank-1 accuracy / mAP compared with the baseline. The proposed method can be generalized to the task of single-modality Re-ID and improves the rank-1 accuracy / mAP from 92.0% / 81.7% to 94.7% / 86.6% on the Market-1501 dataset, from 82.6% / 70.6% to 87.6% / 77.1% on the DukeMTMC-reID dataset.
Hanrong Ye, Hong Liu 0008, Fanyang Meng, Xia Li 0005
IEEE Trans. Image Process.2
2021 When Dictionary Learning Meets Deep Learning: Deep Dictionary Learning and Coding Network for Image Recognition With Limited Data
abstract
We present a new deep dictionary learning and coding network (DDLCN) for image-recognition tasks with limited data. The proposed DDLCN has most of the standard deep learning layers (e.g., input/output, pooling, and fully connected), but the fundamental convolutional layers are replaced by our proposed compound dictionary learning and coding layers. The dictionary learning learns an overcomplete dictionary for input training data. At the deep coding layer, a locality constraint is added to guarantee that the activated dictionary bases are close to each other. Then, the activated dictionary atoms are assembled and passed to the compound dictionary learning and coding layers. In this way, the activated atoms in the first layer can be represented by the deeper atoms in the second dictionary. Intuitively, the second dictionary is designed to learn the fine-grained components shared among the input dictionary atoms; thus, a more informative and discriminative low-level representation of the dictionary atoms can be obtained. We empirically compare DDLCN with several leading dictionary learning methods and deep learning models. Experimental results on five popular data sets show that DDLCN achieves competitive results compared with state-of-the-art methods when the training data are limited. Code is available at https://github.com/Ha0Tang/DDLCN.
Hao Tang 0005, Hong Liu 0008, Wei Xiao 0002, Nicu Sebe
IEEE Trans. Neural Networks Learn. Syst.2
2020 Spatial Pyramid Based Graph Reasoning for Semantic Segmentation
abstract
The convolution operation suffers from a limited receptive filed, while global modeling is fundamental to dense prediction tasks, such as semantic segmentation. In this paper, we apply graph convolution into the semantic segmentation task and propose an improved Laplacian. The graph reasoning is directly performed in the original feature space organized as a spatial pyramid. Different from existing methods, our Laplacian is data-dependent and we introduce an attention diagonal matrix to learn a better distance metric. It gets rid of projecting and re-projecting processes, which makes our proposed method a light-weight module that can be easily plugged into current computer vision architectures. More importantly, performing graph reasoning directly in the feature space retains spatial relationships and makes spatial pyramid possible to explore multiple long-range contextual patterns from different scales. Experiments on Cityscapes, COCO Stuff, PASCAL Context and PASCAL VOC demonstrate the effectiveness of our proposed methods on semantic segmentation. We achieve comparable performance with advantages in computational and memory overhead.
Xia Li 0005, Qijie Zhao, Tiancheng Shen, Zhouchen Lin, Hong Liu 0008
CVPR6
2020 A Fast and Accurate Super-Resolution Network Using Progressive Residual Learning
abstract
Single-image super-resolution (SISR) task has witnessed great strides in the past few years with the development of deep learning. However, most existing studies concentrate on exploiting much deeper super-resolution networks, which are not friendly to the constrained computation resources. In this work, a lightweight network using progressive residual learning for SISR (PRLSR) is proposed to address this issue. Specifically, a progressive residual block (PRB) is designed to progressively downsample deep features for reducing the redundancy and obtaining refined features. Simultaneously, a high-frequency preserving module is proposed to lower the detail loss caused by resolution reduction in PRB. Furthermore, a residual learning-based architecture with learnable weights is utilized to extract multilevel features and adaptively adjust the contribution of residual mapping and identity mapping in residual structure to accelerate convergence. Experimental results on four benchmarks show that our PRLSR achieves superior performance over state-of-the-art methods with a significantly decreased computational cost.
Hong Liu 0008, Zhisheng Lu, Wei Shi 0009, Juanhui Tu
ICASSP1
2020 GFNet: A Lightweight Group Frame Network for Efficient Human Action Recognition
abstract
Human action recognition aims at assigning an action label to a well-segmented video. Recent work using two-stream or 3D convolutional neural networks achieves high recognition rates at the cost of huge computation complexity, memory footprint, and parameters. In this paper, we propose a lightweight neural network called Group Frame Network (GFNet) for human action recognition, which imposes intra-frame spatial information sparsity on spatial dimension in a simple yet effective way. Benefit from two core components, namely Group Temporal Module (GTM) and Group Spatial Module (GSM), GFNet decreases irrelevant motion inside frames and duplicate texture features among frames, which can extract the spatial-temporal information of frames at a minuscule cost. Experimental results on NTU RGB+D dataset and Varying-view RGB-D Action dataset show that our method without any pre-training strategy reaches a reasonable trade-off among computation complexity, parameters and performance, which is more cost-efficient than state-of-the-art methods.
Hong Liu 0008, Lisi Guan
ICASSP1
2020 Position Constraint Loss For Fashion Landmark Estimation
abstract
Fashion landmark estimation aims at locating functional key points of clothes, which has wide potential applications in electronic commerce. However, due to the occlusion and weak outline information, landmark estimation occurs outliers and duplicate detection problems. To alleviate these issues, we propose Position Constraint Loss (PCLoss) to constrain error landmark locations by utilizing the position relationship of landmarks. Specifically, PCLoss adds a regularization term for each landmark to regularize their relative positions, and it can be easily applied to both regression and heatmap based methods without extra computation during inference. Unlike existing approaches that propagate landmark information between feature layers by specific network structures, PCLoss introduces position relations of landmarks in the label space without modifying the network structure. In addition, we leverage the skeleton-like relation of clothing to further strengthen position constraints between landmarks. Extensive experimental results on DeepFashion, FLD and FashionAI demonstrate that our methods can effectively increase the performance of mainstream frameworks by a large margin.
Hong Liu 0008, Meijia Song, Wei Shi 0009, Xia Li 0005
ICASSP1
2020 Spatio-Temporal and Geometry Constrained Network for Automobile Visual Odometry
abstract
Visual odometry (VO) is an essence of vision-based localization and mapping system where existing learning-based approaches utilize CNN and RNN to model camera motion and gain promising results. However, these methods lack full use of the relationship between spatial characteristics and temporal clues, as well as geometry constraints in VO. To overcome these deficiencies, an end-to-end framework that leverages spatio-temporal relevance and geometrical knowledge is proposed. In particular, a spatial response module (SRM) is designed to extract the visual motion features by emphasizing the most interconnected regions while suppressing the irrelevant areas. A module named temporal response module (TRM) is used to regress the camera motion via adopting the optimal motion features. Moreover, a geometry constrained (GC) loss that minimizes the estimated inter-frame pose errors and the accumulated pose errors within a local period is introduced. Actually, the GC loss utilizes adaptive learnable balance factors for balancing losses. Experimental results on KITTI and Malaga datasets demonstrate that the proposed model outperforms state-of-the-art monocular methods.
Hong Liu 0008, Weibo Huang, Guoliang Hua, Fanyang Meng
ICASSP1
2020 Robust Audio-Visual Mandarin Speech Recognition Based On Adaptive Decision Fusion And Tone Features
abstract
Audio-visual speech recognition (AVSR) integrates both audio and visual information to perform automatic speech recognition (ASR), which improves the robustness of human-robot interaction systems especially in noise environments. However, few methods and applications have paid attention to AVSR in tonal languages, in which the linguistic feature can play an important role as well as visual information. In this work, we propose a method for AVSR in Mandarin based on adaptive decision fusion as well as making full use of tone features. Firstly, we introduce tone features calculated by Constant Q trasform (CQT) and put them into a CNN-based audio network together with Mel-Frequency Cepstral Coefficient (MFCC) audio features. Then, the visual features are extracted by Discrete Cosine Transform (DCT) from mouth regions in video frames and modeled by an LSTM-based visual network. Finally, an adaptive decision fusion network combines the outputs from both streams to make final predictions. Experimental results on the PKU-AV2 dataset show that the tone features can significantly improve the robustness of Mandarin speech recognition systems, and the adaptability of the proposed method to various noise environments.
Hong Liu 0008, Zhengyan Chen, Wei Shi 0009
ICIP1
2020 Motion Rectification Network for Unsupervised Learning of Monocular Depth and Camera Motion
abstract
Although unsupervised methods of monocular depth and camera motion estimation have made significant progress, most of them are based on the static scene assumption and may perform poorly in dynamic scenes. In this paper, we propose a novel framework for unsupervised learning of monocular depth and camera motion estimation, which is applicable to dynamic scenes. Firstly, the framework is trained to obtain initial inference results by assuming the scene is static, through minimizing a photometric consistency loss and a 3D transformation consistency loss. Then, the framework is fine-tuned by jointly learning with a motion rectification network (RecNet). Specifically, RecNet is designed to rectify the individual motion of moving objects and generate motion rectified images, enabling the framework to learn accurately in dynamic scenes. Extensive experiments have been done on the KITTI dataset. Results show that our method achieves state-of-the-art performance on both depth prediction and camera motion estimation tasks.
Hong Liu 0008, Guoliang Hua, Weibo Huang
ICIP1
2020 Grouped Temporal Enhancement Module for Human Action Recognition
abstract
Temporal information is a significant cue for recognizing human actions from videos. Different from 2D CNN which can only capture spatial information in an efficient way, 3D CNN is good at capturing both spatial and temporal information at the expense of high computational cost. Beyond both methods, this paper presents a Grouped Temporal Enhancement (GTE) module which even outperforms 3D CNN, meanwhile only needs similar low computational cost as 2D CNN. The GTE module firstly decomposes an input video into spatial and temporal groups along channel dimension, and then uses a learnable temporal shift (LTS) operation for efficient temporal modeling. Finally, a 2D convolution filter is used to enhance the ability of LTS for spatial modeling. Extensive experiments on three benchmark datasets validate the effect of our method.
Hong Liu 0008, Bin Ren 0005, Mengyuan Liu 0001, Runwei Ding
ICIP1
2020 Towards Domain Generalization In Underwater Object Detection
abstract
A General Underwater Object Detector (GUOD) should perform well on most of underwater circumstances. However, with limited underwater dataset, conventional object detection methods suffer from domain shift severely. This paper aims to build a GUOD using small underwater dataset with limited types of water quality. First, we propose a data augmentation method Water Quality Transfer (WQT) to increase domain diversity of the original small dataset. Second, for mining the semantic information from data generated by WQT, Domain Generalization YOLO (DG-YOLO) is proposed, which consists of three parts: YOLOv3, Domain Invariant Module and Invariant Risk Minimization penalty. Finally, experiments on original and synthetic URPC2019 dataset prove that WQT combined with DG-YOLO achieves promising performance of domain generalization in underwater object detection. The source code can be found at https://github.com/mousecpn/DG-YOLO.
Hong Liu 0008, Pinhao Song, Runwei Ding
ICIP1
2020 Efficient High-Resolution High-Level-Semantic Representation Learning for Human Pose Estimation
Hong Liu 0008, Lisi Guan
ICPR1
2020 Robust Audio-Visual Speech Recognition Based on Hybrid Fusion
abstract
The fusion of audio and visual modalities is an important stage of audio-visual speech recognition (AVSR), which is generally approached through feature fusion or decision fusion. Feature fusion can exploit the covariations between features from different modalities effectively, whereas decision fusion shows the robustness of capturing an optimal combination of multimodality. In this work, to take full advantage of the complementarity of the two fusion strategies and address the challenge of inherent ambiguity in noisy environments, we propose a novel hybrid fusion based AVSR method with residual networks and Bidirectional Gated Recurrent Unit (BGRU), which is able to distinguish homophones in both clean and noisy conditions. Specifically, a simple yet effective audio-visual encoder is used to map audio and visual features into a shared latent space to capture more discriminative multi-modal feature and find the internal correlation between spatial-temporal information for different modalities. Furthermore, a decision fusion module is designed to get final predictions in order to robustly utilize the reliability measures of audio-visual information. Finally, we introduce a combined loss, which shows its noise-robustness in learning the joint representation across various modalities. Experimental results on the largest publicly available dataset (LRW) demonstrate the robustness of the proposed method under various noisy conditions.
Hong Liu 0008, Wenhao Li 0002, Bing Yang 0004
ICPR1
2020 A Base-Derivative Framework for Cross-Modality RGB-Infrared Person Re-Identification
abstract
Cross-modality RGB-infrared (RGB-IR) person reidentification (Re-ID) is a challenging research topic due to the heterogeneity of RGB and infrared images. In this paper, we aim to find some auxiliary modalities, which are homologous with the visible or infrared modalities, to help reduce the modality discrepancy caused by heterogeneous images. Accordingly, a new base-derivative framework is proposed, where base refers to the original visible and infrared modalities, and derivative refers to the two auxiliary modalities that are derived from base. In the proposed framework, the double-modality cross-modal learning problem is reformulated as a four-modality one. After that, the images of all the base and derivative modalities are fed into the feature learning network. With the doubled input images, the learned person features become more discriminative. Furthermore, the proposed framework is optimized by the enhanced intra- and cross-modality constraints with the assistance of two derivative modalities. Experimental results on two publicly available datasets SYSU-MM and RegDB show that the proposed method outperforms the other state-of-the-art methods. For instance, we achieve a gain of over 13 % in terms of both Rank- and mAP on RegDB dataset.
Hong Liu 0008, Ziling Miao, Bing Yang 0004, Runwei Ding
ICPR1
2020 3D Audio-Visual Speaker Tracking with A Novel Particle Filter
abstract
3D speaker tracking using co-located audio-visual sensors has received much attention recently. Though various methods have been attempted to this field, it is still challenging to obtain a reliable 3D tracking result since the position of colocated sensors are restricted to a small area. In this paper, a novel particle filter (PF) based method is proposed for 3D audio-visual speaker tracking. Compared with traditional PF based audio-visual speaker tracking method, our 3D audio-visual tracker has two main characteristics. In the prediction stage, we use audio-visual information at current frame to further adjust the direction of the particles after the particle state transition process, which can make the particles more concentrated around the speaker direction. In the update stage, the particle likelihood is calculated by fusing both the visual distance and audiovisual direction information. Specially, the distance likelihood is obtained according to the camera projection model and the adaptively estimated size of speaker face or head, and the direction likelihood is determined by audio-visual particle fitness. In this way, the particle likelihood can better represent the speaker presence probability in 3D space. Experimental results show that the proposed tracker outperforms other methods and provides a favorable speaker tracking performance both in 3D space and on the image plane.
Hong Liu 0008, Yongheng Sun, Yidi Li 0001, Bing Yang 0004
ICPR1
2020 Mutual Alignment between Audiovisual Features for End-to-End Audiovisual Speech Recognition
abstract
Asynchronization issue caused by different types of modalities is one of the major problems in audio visual speech recognition (AVSR) research. However, most AVSR systems merely rely on up sampling of video or down sampling of audio to align audio and visual features, assuming that the feature sequences are aligned frame-by-frame. These pre-processing steps oversimplify the asynchrony relation between acoustic signal and lip motion, lacking flexibility and impairing the performance of the system. Although there are systems modeling the asynchrony between the modalities, sometimes they fail to align speech and video precisely over some even all noisy conditions. In this paper, we propose a mutual feature alignment method for AVSR which can make full use of cross modility information to address the asynchronization issue by introducing Mutual Iterative Attention (MIA) mechanism. Our method can automatically learn an alignment in a mutual way by performing mutual attention iteratively between the audio and visual features, relying on the modified encoder structure of Transformer. Experimental results show that our proposed method obtains absolute improvements up to 20.42% over the audio modality alone depending upon the signal-to-noise-ratio (SNR) level. Better recognition performance can also be achieved comparing with the traditional feature concatenation method under both clean and noisy conditions. It is expectable that our proposed mutual feature alignment method can be easily generalized to other multimodal tasks with semantically correlated information.
Hong Liu 0008, Bing Yang 0004
ICPR1
2020 Audio-Visual Speech Recognition Using A Two-Step Feature Fusion Strategy
abstract
Lip-reading methods and fusion strategy are crucial for audio-visual speech recognition. In recent years, most approaches involve two separate audio and visual streams with early or late fusion strategies. Such a single-stage fusion method may fail to guarantee the integrity and representativeness of fusion information simultaneously. This paper extends a traditional single-stage fusion network to a two-step feature fusion network by adding an audio-visual early feature fusion (AV-EFF) stream to the baseline model. This method can learn the fusion information of different stages, preserving the original features as much as possible and ensuring the independence of different features. Besides, to capture long-range dependencies of video information, a non-local block is added to the feature extraction part of the visual stream (NL-Visual) to obtain the long-term spatio-temporal features. Experimental results on the two largest public datasets in English (LRW) and Mandarin (LRW-1000) demonstrate our method is superior to other state-of-the-art methods.
Hong Liu 0008, Wanlu Xu, Bing Yang 0004
ICPR1
2020 Multi-Scale Cascading Network with Compact Feature Learning for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared person re-identification (RGB-IR Re-ID) aims to matching persons from heterogeneous images captured by visible and thermal cameras, which is of great significance in the surveillance system under poor light conditions. Facing great challenges in complex variances including conventional single-modality and additional inter-modality discrepancies, most of the existing RGB-IR Re-ID methods propose to impose constraints in image level, feature level or a hybrid of both. Despite better performance of hybrid constraints, they are usually implemented with heavy network architecture. As a matter of fact, previous efforts contribute more as pioneering works in new cross-modal Re-ID area while leaving large space for improvement. This can be mainly attributed to: (a) lack of abundant person image pairs from different modalities for training, and (b) scarcity of salient modality-invariant features especially on coarse representations for effective matching. To address these issues, a novel Multi-Scale Part-Aware Cascading framework (MSPAC) is formulated by aggregating multi-scale fine-grained features from part to global in a cascading manner, which results in a unified representation containing rich and enhanced semantic features. Furthermore, a marginal exponential center (MeCen) loss is introduced to jointly eliminate mixed variances from intra- and inter-modal examples. Cross-modality correlations can thus be efficiently explored on salient features for distinctive modality-invariant feature learning. Extensive experiments are conducted to demonstrate that the proposed method outperforms all the state-of-the-art by a large margin.
Can Zhang 0007, Hong Liu 0008, Wei Guo 0006, Mang Ye
ICPR2
2020 Unsupervised Monocular Visual-inertial Odometry Network
abstract
Recently, unsupervised methods for monocular visual odometry (VO), with no need for quantities of expensive labeled ground truth, have attracted much attention. However, these methods are inadequate for long-term odometry task, due to the inherent limitation of only using monocular visual data and the inability to handle the error accumulation problem. By utilizing supplemental low-cost inertial measurements, and exploiting the multi-view geometric constraint and sequential constraint, an unsupervised visual-inertial odometry framework (UnVIO) is proposed in this paper. Our method is able to predict the per-frame depth map, as well as extracting and self-adaptively fusing visual-inertial motion features from image-IMU stream to achieve long-term odometry task. A novel sliding window optimization strategy, which consists of an intra-window and an inter-window optimization, is introduced for overcoming the error accumulation and scale ambiguity problem. The intra-window optimization restrains the geometric inferences within the window through checking the photometric consistency. And the inter-window optimization checks the 3D geometric consistency and trajectory consistency among predictions of separate windows. Extensive experiments have been conducted on KITTI and Malaga datasets to demonstrate the superiority of UnVIO over other state-of-the-art VO / VIO methods. The codes are open-source.
Guoliang Hua, Weibo Huang, Fanyang Meng, Hong Liu 0008
IJCAI5
2020 Lip Graph Assisted Audio-Visual Speech Recognition Using Bidirectional Synchronous Fusion
Hong Liu 0008, Bing Yang 0004
INTERSPEECH1
2020 Part-Based Lipreading for Audio-Visual Speech Recognition
abstract
Lipreading is an important component of audio-visual speech recognition. However, lips are usually modeled as a whole in lipreading, which ignores that each part of lip focuses on different characteristics of mouth and the overall model can not fit each part perfectly. Besides, features based on the whole lip usually vary a lot according to different speakers, which leads that the training databases usually need to contain as much speakers as possible. In this paper, A part-based lipreading (PBL) method is proposed to deal with the mismatch between an overall lip model and the separate parts of lips, also the excessive dependence of models on the speakers in training set. PBL models lips partly and predicts jointly. It employs a uniform partition strategy on convolutional features and generates several part-level sub-results for final prediction. Experiments are performed on a large publicly available dataset (LRW) and part of it (p-LRW, 65 words), in order to simulate the progressive instructions in the working scene of robots. Word accuracy of PBL reaches 82.8% on LRW and 88.9% on p-LRW. Finally, an end-to-end audio-visual speech recognition system using PBL is established and achieves 98.3% word accuracy on LRW.
Ziling Miao, Hong Liu 0008, Bing Yang 0004
SMC2
2020 Identity-sensitive loss guided and instance feature boosted deep embedding for person search
abstract
Person search aims at detecting and re-identifying pedestrians from whole monitoring images, which is vital for intelligent surveillance. However, this task is still challenging due to the extremely few instances per training identity and inherent fine-grained differences among different identities. To this end, this work proposes an identity-sensitive loss guided and instance feature boosted pipeline to extract deep discriminative feature embedding for person search. First, a prior anchor pre-trained network (PAPN) is designed to obtain proper initial state for the whole deep person search training baseline. Second, a new loss function called instance enhancing loss (IEL) is proposed to learn identity-sensitive features by introducing unlabeled identity information. Specifically, the proposed IEL can selectively utilize unlabeled identities with similar appearances to labeled identities to train the person search network. Third, considering the intra-class compactness of features learned by center loss and contextual inter-class relations, two instance boosting strategies (Boosting) are used to learn more discriminative features. Extensive experiments on two benchmark datasets, namely CUHK-SYSU and PRW, demonstrate the effectiveness of our approach.
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001
Neurocomputing2
2020 Attention-guided CNN for image denoising
Chunwei Tian, Yong Xu 0001, Wangmeng Zuo, Lunke Fei, Hong Liu 0008
Neural Networks6
2020 Incomplete Multiview Spectral Clustering With Adaptive Graph Learning
abstract
In this paper, we propose a general framework for incomplete multiview clustering. The proposed method is the first work that exploits the graph learning and spectral clustering techniques to learn the common representation for incomplete multiview clustering. First, owing to the good performance of low-rank representation in discovering the intrinsic subspace structure of data, we adopt it to adaptively construct the graph of each view. Second, a spectral constraint is used to achieve the low-dimensional representation of each view based on the spectral clustering. Third, we further introduce a co-regularization term to learn the common representation of samples for all views, and then use the k -means to partition the data into their respective groups. An efficient iterative algorithm is provided to optimize the model. Experimental results conducted on seven incomplete multiview datasets show that the proposed method achieves the best performance in comparison with some state-of-the-art methods, which proves the effectiveness of the proposed method in incomplete multiview clustering.
Jie Wen 0001, Yong Xu 0001, Hong Liu 0008
IEEE Trans. Cybern.3
2020 Unified Generative Adversarial Networks for Controllable Image-to-Image Translation
abstract
We propose a unified Generative Adversarial Network (GAN) for controllable image-to-image translation, i.e., transferring an image from a source to a target domain guided by controllable structures. In addition to conditioning on a reference image, we show how the model can generate images conditioned on controllable structures, e.g., class labels, object keypoints, human skeletons, and scene semantic maps. The proposed model consists of a single generator and a discriminator taking a conditional image and the target controllable structure as input. In this way, the conditional image can provide appearance information and the controllable structure can provide the structure information for generating the target result. Moreover, our model learns the image-to-image mapping through three novel losses, i.e., color loss, controllable structure guided cycle-consistency loss, and controllable structure guided self-content preserving loss. Also, we present the Fr´echet ResNet Distance (FRD) to evaluate the quality of the generated images. Experiments on two challenging image translation tasks, i.e., hand gesture-to-gesture translation and cross-view image translation, show that our model generates convincing results, and significantly outperforms other state-of-the-art methods on both tasks. Meanwhile, the proposed framework is a unified solution, thus it can be applied to solving other controllable structure guided image translation tasks such as landmark guided facial expression translation and keypoint guided person image generation. To the best of our knowledge, we are the first to make one GAN framework work on all such controllable structure guided image translation tasks. Code is available at https://github.com/Ha0Tang/GestureGAN.
Hao Tang 0005, Hong Liu 0008, Nicu Sebe
IEEE Trans. Image Process.2
2020 An Online Initialization and Self-Calibration Method for Stereo Visual-Inertial Odometry
abstract
Most online initialization and self-calibration methods for visual-inertial odometry (VIO) are only able to estimate the extrinsic parameters (orientation and translation) between one camera and inertial measurement unit (IMU) pair. They are not applicable to stereo VIO where both camera-IMU and camera-camera pairs exist. In this article, we address the issue by taking advantage of the geometric constraints among the multiple sensors. An online method is proposed to estimate the initial values of velocity, gravity, IMU biases, and simultaneously calibrate the extrinsic parameters of camera-camera and camera-IMU pairs for bootstrapping a smoothing-based stereo VIO system. The method includes a three-step process to incrementally solve several linear equations in a coarse-to-fine manner. It back-propagates historically estimated results to update weight factors and remove outliers, and employs a convergence criterion to monitor and terminate the process. It also includes an optional global optimization for further refinement. The method is evaluated in terms of accuracy, robustness, convergence, consistency, and tunable parameters using both simulated and public datasets. Experimental results show that the proposed method can accurately estimate the initial values and the extrinsic parameters.
Weibo Huang, Hong Liu 0008, Weiwei Wan
IEEE Trans. Robotics2
2019 Unified Embedding Alignment with Missing Views Inferring for Incomplete Multi-View Clustering
abstract
Multi-view clustering aims to partition data collected from diverse sources based on the assumption that all views are complete. However, such prior assumption is hardly satisfied in many real-world applications, resulting in the incomplete multi-view learning problem. The existing attempts on this problem still have the following limitations: 1) the underlying semantic information of the missing views is commonly ignored; 2) The local structure of data is not well explored; 3) The importance of different views is not effectively evaluated. To address these issues, this paper proposes a Unified Embedding Alignment Framework (UEAF) for robust incomplete multi-view clustering. In particular, a locality-preserved reconstruction term is introduced to infer the missing views such that all views can be naturally aligned. A consensus graph is adaptively learned and embedded via the reverse graph regularization to guarantee the common local structure of multiple views and in turn can further align the incomplete views and inferred views. Moreover, an adaptive weighting strategy is designed to capture the importance of different views. Extensive experimental results show that the proposed method can significantly improve the clustering performance in comparison with some state-of-the-art methods.
Jie Wen 0001, Zheng Zhang 0006, Yong Xu 0001, Bob Zhang 0001, Lunke Fei, Hong Liu 0008
AAAI6
2019 A Weight-shared Dual-branch Convolutional Neural Network for Unsupervised Dense Depth Prediction and Camera Motion Estimation
abstract
Convolutional Neural Network (CNN) can be used to indiscriminately predict dense depth and camera motion from images, however, ignoring the relationship between depth map and camera motion increases the computational burden to label the datasets and limits the accuracy of the results. In this paper, an end-to-end unsupervised dual-branch CNN is proposed to predict a pixel-wise depth map and simultaneously estimate camera pose. In particular, a weight sharing strategy for two branches is designed to increase the connection between depth map and camera motion. Besides, to reduce the impact of photometric noise, the intermediate feature maps are utilized to compute feature errors. Experimental results on the KITTI datasets demonstrate that our method achieves better performance on dense map prediction and camera pose estimation comparing with the state-of-the-art approaches.
Hong Liu 0008, Yaofeng Dong, Weibo Huang
ICASSP1
2019 Expectation-Maximization Attention Networks for Semantic Segmentation
abstract
Self-attention mechanism has been widely used for various tasks. It is designed to compute the representation of each position by a weighted sum of the features at all positions. Thus, it can capture long-range relations for computer vision tasks. However, it is computationally consuming. Since the attention maps are computed w.r.t all other positions. In this paper, we formulate the attention mechanism into an expectation-maximization manner and iteratively estimate a much more compact set of bases upon which the attention maps are computed. By a weighted summation upon these bases, the resulting representation is low-rank and deprecates noisy information from the input. The proposed Expectation-Maximization Attention (EMA) module is robust to the variance of input and is also friendly in memory and computation. Moreover, we set up the bases maintenance and normalization methods to stabilize its training procedure. We conduct extensive experiments on popular semantic segmentation benchmarks including PASCAL VOC, PASCAL Context, and COCO Stuff, on which we set new records1.
Xia Li 0005, Zhisheng Zhong, Jianlong Wu, Zhouchen Lin, Hong Liu 0008
ICCV6
2019 Self-Refining Deep Symmetry Enhanced Network for Rain Removal
abstract
Rain removal aims to remove the rain streaks on rain images. The state-of-the-art methods are mostly based on Convolutional Neural Network (CNN). However, as CNN is not equivariant to object rotation, these methods are unsuitable for dealing with the tilted rain streaks. To tackle this problem, we propose Deep Symmetry Enhanced Network (DSEN) that is able to explicitly extract the rotation equivariant features from rain images. In addition, we design a self-refining mechanism to remove the accumulated rain streaks in a coarse-to-fine manner. This mechanism reuses DSEN with a novel information link which passes the gradient flow to the higher stages. Extensive experiments on both synthetic and real-world rain images show that our self-refining DSEN yields the top performance.
Hong Liu 0008, Hanrong Ye, Xia Li 0005, Wei Shi 0009, Mengyuan Liu 0001, Qianru Sun
ICIP1
2019 3D Audio-Visual Speaker Tracking with A Two-Layer Particle Filter
abstract
Audio-visual speaker tracking in 3D space is a challenging problem. Although the classical particle filter based methods have shown effectiveness in audio-visual speaker tracking, the performance degrades considerably when the measurements are disturbed by noise. To this end, a novel two-layer particle filter is proposed for 3D audio-visual speaker tracking. Firstly, two groups of particles, which are generated from the audio and video streams respectively, are propagated independently in the audio layer and visual layer. Then, the audio and visual likelihoods are combined in an adaptive sigmoid function, which can adjust particle weights according to the confidence of two modalities. Finally, an optimal particle set selected from two groups of particles is proposed to determine the speaker position and reset the particle positions in the next frame. Experiments on AV16.3 database show that our method outperforms the trackers using individual modalities and the existing approaches in the 3D space and on the image plane.
Hong Liu 0008, Yidi Li 0001, Bing Yang 0004
ICIP1
2019 Fast and robust dynamic hand gesture recognition via key frames extraction and feature fusion
Hao Tang 0005, Hong Liu 0008, Wei Xiao 0002, Nicu Sebe
Neurocomputing2
2019 Multiple Sound Source Counting and Localization Based on TF-Wise Spatial Spectrum Clustering
abstract
This paper addresses the problem of multiple sound source counting and localization in adverse acoustic environments, using microphone array recordings. The proposed time-frequency (TF) wise spatial spectrum clustering based method contains two stages. First, given the received sensor signals, the spatial correlation matrix is computed and denoised in the TF domain. The TF-wise spatial spectrum is estimated based on the signal subspace information, and further enhanced by an exponential transform, which can increase the reliability of the source presence possibility reflected by spatial spectrum. Second, to jointly count and localize sound sources, the enhanced TF-wise spatial spectra are divided into several clusters with each cluster corresponding to one source. Sources are successively detected by searching the significant peaks of the remaining global spatial spectrum, which is formed using unassigned spatial spectra. After each new source detection, spatial spectra are reassigned to detected sources according to the dominance association between them. The interaction between sources is reduced by iteratively performing new source detection and spatial spectrum assignment. Experiments on both simulated data and real-world data demonstrate the superiority of the proposed method for multiple sound source counting and localization in the environment with different levels of noise and reverberation.
Bing Yang 0004, Hong Liu 0008, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Sample Fusion Network: An End-to-End Data Augmentation Network for Skeleton-Based Human Action Recognition
abstract
Data augmentation is a widely used technique for enhancing the generalization ability of deep neural networks for skeleton-based human action recognition (HAR) tasks. Most existing data augmentation methods generate new samples by means of handcrafted transforms. However, these methods often cannot be trained and then are discarded during testing because of the lack of learnable parameters. To solve those problems, a novel type of data augmentation network called a sample fusion network (SFN) is proposed. Instead of using handcrafted transforms, an SFN generates new samples via a long short-term memory (LSTM) autoencoder (AE) network. Therefore, an SFN and HAR network can be cascaded together to form a combined network that can be trained in an end-to-end manner. Moreover, an adaptive weighting strategy is employed to improve the complementarity between a sample and the new sample generated from it by an SFN, thus allowing the SFN to more efficiently improve the performance of the HAR network during testing. The experimental results on various datasets verify that the proposed method outperforms state-of-the-art data augmentation methods. More importantly, the proposed SFN architecture is a general framework that can be integrated with various types of networks for HAR. For example, when a baseline HAR model with three LSTM layers and one fully connected (FC) layer was used, the classification accuracy was increased from 79.53% to 90.75% on the NTU RGB+D dataset using a cross-view protocol, thus outperforming most other methods.
Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Juanhui Tu, Mengyuan Liu 0001
IEEE Trans. Image Process.2
2018 Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation
abstract
Recent works have shown the benefit of integrating Conditional Random Fields (CRFs) models into deep architectures for improving pixel-level prediction tasks. Following this line of research, in this paper we introduce a novel approach for monocular depth estimation. Similarly to previous works, our method employs a continuous CRF to fuse multi-scale information derived from different layers of a front-end Convolutional Neural Network (CNN). Differently from past works, our approach benefits from a structured attention model which automatically regulates the amount of information transferred between corresponding features at different scales. Importantly, the proposed attention model is seamlessly integrated into the CRF, allowing end-to-end training of the entire architecture. Our extensive experimental evaluation demonstrates the effectiveness of the proposed method which is competitive with previous methods on the KITTI benchmark and outperforms the state of the art on the NYU Depth V2 dataset.
Dan Xu 0002, Wei Wang 0108, Hao Tang 0005, Hong Liu 0008, Nicu Sebe, Elisa Ricci 0001
CVPR4
2018 Recurrent Squeeze-and-Excitation Context Aggregation Net for Single Image Deraining
Xia Li 0005, Jianlong Wu, Zhouchen Lin, Hong Liu 0008, Hongbin Zha
ECCV (7)4
2018 A Discriminatively Learned Feature Embedding Based on Multi-Loss Fusion For Person Search
abstract
Person search is a challenging task that requires to address pedestrian detection and person re- identification simultaneously. Though significant progress has been made in detection and re-identification respectively, the similar appearances of persons, pedestrian misdetections and false alarms still have adverse effects on person search. To this end, an improved end-to-end person search network with multi -loss is proposed to jointly optimize detection and re-identification. Firstly, a pre-trained network is designed to obtain proper initial state for the whole training network. Then, to enhance the person search model, an improved online instance matching (IOIM) loss is proposed by hardening the distribution of labeled identities and softening the distribution of unlabeled identities. Finally, considering the intra-class compactness of features learned by center loss, the IOIM loss is combined with center loss by the proposed multi-loss fusion strategy, which can learn more discriminative feature embeddings. Experimental results on two challenging datasets CUHK-SYSU and PRW demonstrate our approach significantly outperforms the state-of-the-arts.
Hong Liu 0008, Wei Shi 0009, Weipeng Huang, Qiao Guan
ICASSP1
2018 Learning Explicit Shape and Motion Evolution Maps for Skeleton-Based Human Action Recognition
abstract
Human action recognition based on skeleton sequences has wide applications in human-computer interaction and intelligent surveillance. Although previous methods have successfully applied Long Short-Term Memory(LSTM) networks to model shape evolution of human actions, it still remains a problem to efficiently recognize actions, especially for similar actions from sequential data due to the lack of the details of motion. To solve this problem, this paper presents an improved LSTM-based network to jointly learn explicit long-term shape evolution maps (SEM) and motion evolution maps (MEM). Firstly, human actions are represented as compact SEM and MEM, which mutually compensate. Secondly, these maps are jointly learned by deep LSTM networks to explore high-level temporal dependencies. Then, a weighted aggregate layer (WAL) is designed to aggregate outputs of L-STM networks cross different temporal stages. Finally, deep features of shape and motion are combined by decision level fusion. Experimental results on the currently largest NTU RGB+D dataset and public SmartHome dataset verify that our method significantly outperforms the state-of-the-arts.
Hong Liu 0008, Juanhui Tu, Runwei Ding
ICASSP1
2018 An End-To-End Siamese Convolutional Neural Network for Loop Closure Detection in Visual Slam System
abstract
Loop closure detection is essential and important in visual simultaneous localization and mapping (SLAM) systems. Most existing methods typically utilize a separate feature extraction part and a similarity metric part. Compared to these methods, an end-to-end network is proposed in this paper to jointly optimize the two parts in a unified framework for further enhancing the interworking between these two parts. First, a two-branch siamese network is designed to learn respective features for each scene of an image pair. Then a hierarchical weighted distance (HWD) layer is proposed to fuse the multi-scale features of each convolutional module and calculate the distance between the image pair. Finally, by using the contrastive loss in the training process, the effective feature representation and similarity metric can be learned simultaneously. Experiments on several open datasets illustrate the superior performance of our approach and demonstrate that the end-to-end network is feasible to conduct the loop closure detection in real time and provides an implementable method for visual SLAM systems.
Hong Liu 0008, Weipeng Huang, Wei Shi 0009
ICASSP1
2018 Audio-Visual Keyword Spotting Based on Multidimensional Convolutional Neural Network
abstract
The fusion of audio and visual information is one of the most promising solutions for reliable keyword spotting (KWS), particularly when audio is corrupted by noise. KWS aims to detect a specific word in an audio stream, which still remains a challenging problem under noisy environments. In this paper, an audio-visual neural network based on multidimensional convolutional neural network (MCNN) is proposed to perform audio-visual KWS. Firstly, the log mel-spectrogram and lip area sequence are extracted, respectively, from the audio and visual streams, and are taken as the input of the audio-visual neural network. Then, an audio-visual neural network based on MCNN consisting of 2D CNN and 3D CNN is used to model the time-frequency feature of the log mel-spectrogram and the spatiotemporal feature of the lip area sequence, respectively. Finally, the outputs of the audio and visual networks are combined for KWS through decision fusion. Experimental results on the PKU-AV database under complex acoustic conditions demonstrate that the proposed method achieves preferable performance compared to other state-of-the-art methods.
Runwei Ding, Hong Liu 0008
ICIP3
2018 Instance Enhancing Loss: Deep Identity-Sensitive Feature Embedding for Person Search
abstract
Person search, which is vital for intelligent surveillance, aims at detecting and re-identifying pedestrians from whole monitoring images. However, due to the inaccurate pedestrian detections and extremely few instances per training identity, it remains challenging to learn discriminative representations only by labeled identities for person search. To this end, this paper proposes a novel loss function called instance enhancing loss (IEL) to learn deep identity-sensitive features by introducing unlabeled identity information. Specifically, the proposed IEL can selectively annotate unlabeled identities with similar appearances to labeled identities, and utilize these unlabeled identities in conjunction with labeled identities to train the person search network. The amount of unlabeled identities used as labeled instances can be quantitatively adjusted. Moreover, the proposed IEL is trainable and easy to optimize by back propagation algorithms. Extensive experiments on two benchmark datasets, namely CUHK-SYSU and PRW, show that our method outperforms state-of-the-arts for person search.
Wei Shi 0009, Hong Liu 0008, Fanyang Meng, Weipeng Huang
ICIP2
2018 Spatial-Temporal Data Augmentation Based on LSTM Autoencoder Network for Skeleton-Based Human Action Recognition
abstract
Data augmentation is known to be of crucial importance for the generalization of RNN-based methods of skeleton-based human action recognition. Traditional data augmentation methods artificially adopt various transformations merely in spatial domain, which lack effective temporal representation. This paper extends traditional Long Short-Term Memory (LSTM) and presents a novel LSTM autoencoder network (LSTM-AE) for spatial-temporal data augmentation. In the LSTM-AE, the LSTM network preserves the temporal information of skeleton sequences, and the autoencoder architecture can automatically eliminate irrelevant and redundant information. Meanwhile, a regularized cross-entropy loss is defined to guide the LSTM-AE to learn more suitable representations of skeleton data. Experimental results on the currently largest NTU RGB+D dataset and public SmartHome dataset verify that the proposed model outperforms the state-of-the-art methods, and can be integrated with most of the RNN-based action recognition models easily.
Juanhui Tu, Hong Liu 0008, Fanyang Meng, Mengyuan Liu 0001, Runwei Ding
ICIP2
2018 Hierarchical Dropped Convolutional Neural Network for Speed Insensitive Human Action Recognition
abstract
Human action recognition using skeleton data has lots of potential applications in content-based action retrieval and intelligent surveillance, with wide usage of depth sensors and robust skeleton estimation algorithms. Previous methods describe spatial temporal skeleton joints as a compact color image and then use Convolutional Neural Network (CNN) to extract more discriminative deep features. However, these methods ignore the effect of speed variation, which is a common phenomenon and can bring severe intra-varieties to same types of actions. To solve this problem, this paper presents a novel hierarchical dropped CNN architecture, which is constructed in two stages. Dropped CNN (d-CNN) is firstly developed to extract deep features from a probabilistic speed insensitive color image. This image expresses both spatial distributions and temporal evolutions of skeleton joints meanwhile avoids the effect of speed variations. To enhance the temporal discriminative power, we extend d-CNN to a hierarchical structure (h-CNN), where multiple scales of temporal information are encoded. Extensive experiments on benchmark MSRC-12 dataset and the largest NTU RGB+D dataset verify the effectiveness and robustness of the proposed method.
Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Mengyuan Liu 0001, Wei Liu 0065
ICME2
2018 Skeleton-Based Human Action Recognition Using Spatial Temporal 3D Convolutional Neural Networks
abstract
It remains a challenge to extract spatial-temporal information from skeleton sequences for 3D human action recognition. Although most recent action recognition methods based on Recurrent Neural Networks (RNN) have achieved outstanding performance, one of the shortcomings of these methods is the tendency to overemphasize the temporal information. Since 3D Convolutional Neural Networks(3D CNN) can simultaneously learn features from both spatial and temporal dimensions through capturing correlations among three-dimensional signals, this paper proposes a novel two-stream model using 3D CNN. To our best knowledge, this is the first attempt to use 3D CNN in the field of skeleton-based action recognition. Our method consists of three stages. First, skeleton joints are mapped into a 3D coordinate space to encode the spatial and temporal information. Second, 3D CNN models are separately employed to extract deep features from both spatial and temporal stream. Third, to enhance the ability of discriminative features to capture global relationships, we extend each stream into multi-temporal version. Extensive experiments on the large-scale NTU RGB-D dataset and the public SmartHome dataset demonstrate that our method outperforms most of RNN-based methods, which verify the complementary property between spatial and temporal information and the robustness to noise.
Juanhui Tu, Hong Liu 0008
ICME3
2018 Online Initialization and Automatic Camera-IMU Extrinsic Calibration for Monocular Visual-Inertial SLAM
abstract
Most of the existing monocular visual-inertial SLAM techniques assume that the camera-IMU extrinsic parameters are known, therefore these methods merely estimate the initial values of velocity, visual scale, gravity, biases of gyroscope and accelerometer in the initialization stage. However, it's usually a professional work to carefully calibrate the extrinsic parameters, and it is required to repeat this work once the mechanical configuration of the sensor suite changes slightly. To tackle this problem, we propose an online initialization method to automatically estimate the initial values and the extrinsic parameters without knowing the mechanical configuration. The biases of gyroscope and accelerometer are considered in our method, and a convergence criteria for both orientation and translation calibration is introduced to identify the convergence and to terminate the initialization procedure. In the three processes of our method, an iterative strategy is firstly introduced to iteratively estimate the gyroscope bias and the extrinsic orientation. Secondly, the scale factor, gravity, and extrinsic translation are approximately estimated without considering the accelerometer bias. Finally, these values are further optimized by a refinement algorithm in which the accelerometer bias and the gravitational magnitude are taken into account. Extensive experimental results show that our method achieves competitive accuracy compared with the state-of-the-art with less calculation.
Weibo Huang, Hong Liu 0008
ICRA2
2018 Multiple Concurrent Sound Source Tracking Based on Observation-Guided Adaptive Particle Filter
Hong Liu 0008, Haipeng Lan, Bing Yang 0004
INTERSPEECH1
2018 3D Action Recognition Using Multiscale Energy-Based Global Ternary Image
abstract
This paper presents an effective multiscale energy-based global ternary image (E-GTI) representation for action recognition from depth sequences. The unique property of our representation is that it takes the spatiotemporal discrimination and action speed variations into account, intending to solve the problems of distinguishing similar actions and identifying the actions with different speeds in one goal. The entire method is carried out in two stages. In the first stage, consecutive depth frames are used to generate global ternary image (GTI) features, which implicitly capture both inter-frame motion regions and motion directions. Specifically, each pixel in the GTI represents one of three possible states, namely, positive, negative, and neutral, which indicate the increased, decreased, and same depth values, respectively. To cope with speed variations in actions, energy-based sampling method is utilized, leading to multiscale E-GTI features, where the multiscale scheme can efficiently capture the temporal relationships among frames. In the second stage, all the E-GTI features are transformed by Radon transform (RT) as robust descriptors, which are aggregated by the bag-of-visual-words model as a compact representation. Extensive experiments on benchmark data sets show that our representation outperforms state-of-the-art approaches, since it captures discriminating spatiotemporal information of actions. Due to the merits of energy-based sampling and RT methods, our representation shows robustness to speed variations, depth noise, and partial occlusions.
Mengyuan Liu 0001, Hong Liu 0008, Chen Chen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Robust 3D Action Recognition Through Sampling Local Appearances and Global Distributions
abstract
Three-dimensional (3-D) action recognition has broad applications in human-computer interaction and intelligent surveillance. However, recognizing similar actions remains challenging since previous literature fails to capture motion and shape cues effectively from noisy depth data. In this paper, we propose a novel two-layer Bag-of-Visual-Words (BoVW) model, which suppresses the noise disturbances and jointly encodes both motion and shape cues. First, background clutter is removed by a background modeling method that is designed for depth data. Then, motion and shape cues are jointly used to generate robust and distinctive spatial-temporal interest points (STIPs): motion-based STIPs and shape-based STIPs. In the first layer of our model, a multiscale 3-D local steering kernel descriptor is proposed to describe local appearances of cuboids around motion-based STIPs. In the second layer, a spatial-temporal vector descriptor is proposed to describe the spatial-temporal distributions of shape-based STIPs. Using the BoVW model, motion and shape cues are combined to form a fused action representation. Our model performs favorably compared with common STIP detection and description methods. Thorough experiments verify that our model is effective in distinguishing similar actions and robust to background clutter, partial occlusions and pepper noise.
Mengyuan Liu 0001, Hong Liu 0008, Chen Chen 0001
IEEE Trans. Multim.2
2017 Human action recognition using Adaptive Hierarchical Depth Motion Maps and Gabor filter
abstract
Depth motion maps (DMMs) have shown effectiveness in human action recognition, however, they lose the temporal information and suffer from intra-class variations caused by action speed variations. To address these challenges, we propose a novel method for human action recognition. Firstly, Adaptive Hierarchical Depth Motion Maps (AH-DMMs) are calculated over temporal hierarchical windows of video sequences to capture the temporal information. Moreover, adaptive windows and steps are employed to ensure that AH-DMMs are robust to motion speed variations. Then, Gabor filter is adopted to encode the texture information of AH-DMMs, generating compact and discriminative action representations. Finally, the representations serve as the input of collaborative representation classifier (CRC). Experimental results on public benchmark MSRAction3D dataset and DHA dataset demonstrate the superiority of the proposed method over the state-of-the-art depth-based action recognition approaches.
Hong Liu 0008, Qinqin He
ICASSP1
2017 A novel re-tracking strategy for monocular SLAM
abstract
Tracking failure is an inevitable event in real-time monocular simultaneous localization and mapping (SLAM) system. Relocalization procedure that focuses on exploring a visited place to relocalize a camera is usually used to handle this problem. However, this strategy abandons the poses of the frames captured after tracking failure, resulting in losing a part of trajectory with respect to the ground truth. Therefore, a re-tracking strategy (RTS) is proposed to estimate these abandoned frames. An automatic local initialization is activated to initialize a new tracking process when tracking fails. Trajectory correction and fusion are employed when a loop closure is detected. When the last frame of the sequence is detected, the trajectories that cannot loop with the original trajectory should be culled by trajectory culling procedure. Experimental results on two challenging datasets, TUM RGB-D and NewCollege, indicate that the proposed method achieves low root mean square error (RMSE) and high trajectory completeness rate (TCR), especially for rapid moving camera.
Hong Liu 0008, Weibo Huang
ICASSP1
2017 Multiple sound source localization based on TDOA clustering and multi-path matching pursuit
abstract
Multiple sound source localization in wireless acoustic sensor networks (WASNs) is a challenging problem. Although compressive sensing based methods have shown effectiveness in uncorrelated sources localization, their performance degrades significantly when they are used to locate multiple speech sources. To this end, we propose a multiple sound source localization method based on the time difference of arrival (TDOA) clustering and the multi-path matching pursuit algorithm. First, TDOAs are calculated locally in time-frequency (TF) bins of sensor recordings. Then, the TDOAs are clustered after utilizing outlier rejection to remove erroneous estimations. Finally, a multi-path matching pursuit algorithm is proposed to solve a sparse localization model for localizing multiple sound sources. Experimental results show that the proposed method yields good performance for multiple sound source localization, especially in strong noisy scenarios.
Hong Liu 0008, Bing Yang 0004
ICASSP1
2017 Fusing shape and motion matrices for view invariant action recognition using 3D skeletons
abstract
Action recognition under arbitrary views remains a challenge, since view variations bring severe motion and appearance changes which increase the ambiguities among same types of actions. To solve this problem, we propose a new method to effectively capture view invariant shape and motion cues. This method contains three main stages. First, we compute distances among pairwise skeleton joints to form a distance matrix for each skeleton. Second, shape matrices (SMs) and motion matrices (MMs) are formulated to describe shape and motion cues between pairwise distance matrices, respectively. Third, Fisher Vector and Linear Discriminant Analysis (L-DA) are adopted to encode SMs and MMs as low dimension and high discriminative representations, which are further fused to generate final action representation. Experimental results on benchmark UTKinect-Action dataset show that our method achieves better results than methods designed for view invariant action recognition task. Additionally, we collect a SmartHome dataset, on which the robustness of our method to noisy skeleton data is verified.
Qinqin He, Hong Liu 0008
ICIP3
2017 A bidirectional adaptive bandwidth mean shift strategy for clustering
abstract
The bandwidth of a kernel function is a crucial parameter in the mean shift algorithm. This paper proposes a novel adaptive bandwidth strategy which contains three main contributions. (1) The differences among different adaptive bandwidth are analyzed. (2) A new mean shift vector based on bidirectional adaptive bandwidth is defined, which combines the advantages of different adaptive bandwidth strategies. (3) A bidirectional adaptive bandwidth mean shift (BAMS) strategy is proposed to improve the ability to escape from the local maximum density. Compared with contemporary adaptive bandwidth mean shift strategies, experiments demonstrate the effectiveness of the proposed strategy.
Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Liu Wei, Jihong Pei
ICIP2
2017 Time-ordered spatial-temporal interest points for human action classification
abstract
Human action classification, which is vital for content-based video retrieval and human-machine interaction, finds problem in distinguishing similar actions. Previous works typically detect spatial-temporal interest points (STIPs) from action sequences and then adopt bag-of-visual words (BoVW) model to describe actions as numerical statistics of STIPs. Despite the robustness of BoVW, this model ignores the spatial-temporal layout of STIPs, leading to misclassification among different types of actions with similar numerical statistics of STIPs. Motivated by this, a time-ordered feature is designed to describe the temporal distribution of STIPs, which contains complementary structural information to traditional BoVW model. Moreover, a temporal refinement method is used to eliminate intra-variations among time-ordered features caused by performers' habits. Then a time-ordered BoVW model is built to represent actions, which encodes both numerical statistics and temporal distribution of STIPs. Extensive experiments on three challenging datasets, i.e., KTH, Rochster and UT-Interaction, validate the effectiveness of our method in distinguishing similar actions.
Chen Chen 0001, Hong Liu 0008
ICME3
2017 Learning informative pairwise joints with energy-based temporal pyramid for 3D action recognition
abstract
This paper presents an effective local spatial-temporal descriptor for action recognition from skeleton sequences. The unique property of our descriptor is that it takes the spatial-temporal discrimination and action speed variations into account, intending to solve the problems of distinguishing similar actions and identifying actions with different speeds in one goal. The entire algorithm consists of two stages. First, a frame selection method is used to remove noisy skeletons for a given skeleton sequence. From the selected skeletons, skeleton joints are mapped to a high dimensional space, where each point refers to kinematics, time label and joint label of a skeleton joint. To encode relative relationships among joints, pairwise points from the space are then jointly mapped to a new space, where each point encodes the relative relationships of skeleton joints. Second, Fisher Vector (FV) is employed to encode all points from the new space as a compact feature representation. To cope with speed variations in actions, an energy-based temporal pyramid is applied to form a multi-temporal FV representation, which is fed into a kernel-based extreme learning machine classifier for recognition. Extensive experiments on benchmark datasets consistently show that our method outperforms state-of-the-art approaches for skeleton-based action recognition.
Chen Chen 0001, Hong Liu 0008
ICME3
2017 3D action recognition using data visualization and convolutional neural networks
abstract
It remains a challenge to efficiently represent spatial-temporal data for 3D action recognition. To solve this problem, this paper presents a new skeleton-based action representation using data visualization and convolutional neural networks, which contains four main stages. First, skeletons from an action sequence are mapped as a set of five dimensional points, containing three dimensions of location, one dimension of time label and one dimension of joint label. Second, these points are encoded as a series of color images, by visualizing points as RGB pixels. Third, convolutional neural networks are adopted to extract deep features from color images. Finally, action class score is calculated by fusing selected deep features. Extensive experiments on three benchmark datasets show that our method achieves state-of-the-art results.
Chen Chen 0001, Hong Liu 0008
ICME3
2017 Multiple Sound Source Counting and Localization Based on Spatial Principal Eigenvector
Bing Yang 0004, Hong Liu 0008
INTERSPEECH2
2017 How do you smile? Towards a comprehensive smile analysis system
Hong Liu 0008, Chao Xu 0006, Yuan Gao 0008, Xuewu Zhang 0003
Neurocomputing2
2017 Spontaneous versus posed smile recognition via region-specific texture descriptor and geometric facial dynamics
abstract
As a typical biometric cue with great diversities, smile is a fairly influential signal in social interaction, which reveals the emotional feeling and inner state of a person. Spontaneous and posed smiles initiated by different brain systems have differences in both morphology and dynamics. Distinguishing the two types of smiles remains challenging as discriminative subtle changes need to be captured, which are also uneasily observed by human eyes. Most previous related works about spontaneous versus posed smile recognition concentrate on extracting geometric features while appearance features are not fully used, leading to the loss of texture information. In this paper, we propose a region-specific texture descriptor to represent local pattern changes of different facial regions and compensate for limitations of geometric features. The temporal phase of each facial region is divided by calculating the intensity of the corresponding facial region rather than the intensity of only the mouth region. A mid-level fusion strategy of support vector machine is employed to combine the two feature types. Experimental results show that both our proposed appearance representation and its combination with geometry-based facial dynamics achieve favorable performances on four baseline databases: BBC, SPOS, MMI, and UvA-NEMO.
Hong Liu 0008, Xuewu Zhang 0003, Yuan Gao 0008
Frontiers Inf. Technol. Electron. Eng.2
2017 Enhanced skeleton visualization for view invariant human action recognition
Hong Liu 0008, Chen Chen 0001
Pattern Recognit.2
2017 Online growing neural gas for anomaly detection in changing surveillance scenes
Qianru Sun, Hong Liu 0008, Tatsuya Harada
Pattern Recognit.2
2017 Binaural Sound Localization Based on Reverberation Weighting and Generalized Parametric Mapping
abstract
Binaural sound source localization is an important technique for speech enhancement, video conferencing, and human-robot interaction, etc. However, in realistic scenarios, the reverberation and environmental noise would degrade the precision of sound direction estimation. Therefore, reliable sound localization is essential to practical applications. To deal with these disturbances, this paper presents a novel binaural sound source localization approach based on reverberation weighting and generalized parametric mapping. First, the reverberation weighting as a preprocessing stage, is used to separately suppress the early and late reverberation, while preserving interaural cues. Then, two binaural cues, i.e., interaural time and intensity differences, are extracted from the frequency-domain representations of dereverberated binaural signals for the online localization. Their corresponding templates are established using the training data. Furthermore, the generalized parametric mapping is proposed to build a generalized parametric model for describing relationships between azimuth and binaural cues analytically. Finally, a two-step sound localization process is introduced to refine azimuth estimation based on the generalized parametric model and template matching. Experiments in both simulated and real scenarios validate that the proposed method can achieve better localization performance compared to state-of-the-art methods.
Hong Liu 0008, Jie Zhang 0042, Xiaofei Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Energy-Based Global Ternary Image for Action Recognition Using Sole Depth Sequences
abstract
In order to efficiently recognize actions from depth sequences, we propose a novel feature, called Global Ternary Image (GTI), which implicitly encodes both motion regions and motion directions between consecutive depth frames via recording the changes of depth pixels. In this study, each pixel in GTI indicates one of the three possible states, namely positive, negative and neutral, which represents increased, decreased and same depth values, respectively. Since GTI is sensitive to the subject's speed, we obtain energy-based GTI (E-GTI) by extracting GTI from pairwise depth frames with equal motion energy. To involve temporal information among depth frames, we extract E-GTI using multiple settings of motion energy. Here, the noise can be effectively suppressed by describing E-GTIs using the Radon Transform (RT). The 3D action representation is formed as a result of feeding the hierarchical combination of RTs to the Bag of Visual Words model (BoVW). From the extensive experiments on four benchmark datasets, namely MSRAction3D, DHA, MSRGesture3D and SKIG, it is evident that the hierarchical E-GTI outperforms the existing methods in 3D action recognition. We tested our proposed approach on extended MSRAction3D dataset to further investigate and verify its robustness against partial occlusions, noise and speed.
Hong Liu 0008, Chen Chen 0001, Maryam Najafian
3DV2
2016 3D Action Recognition Using Multi-Temporal Depth Motion Maps and Fisher Vector
Chen Chen 0001, Baochang Zhang 0001, Jungong Han, Junjun Jiang, Hong Liu 0008
IJCAI6
2016 A Novel Feature Matching Strategy for Large Scale Image Retrieval
Hao Tang 0005, Hong Liu 0008
IJCAI2
2016 Multi-Channel Linear Prediction Based on Binaural Coherence for Speech Dereverberation
Hong Liu 0008
INTERSPEECH1
2016 Probabilistic binaural multiple sources localization based on time-delay compensation estimator and clustering analysis
abstract
Sound source localization (SSL) is an essential technique in many applications, such as robot audition, human-robot interaction and speech capturing. However, SSL from a binaural input is still a challenging problem, particularly when multiple sources are active simultaneously. In this work, we propose a multi-sources localization framework based on the time-delay compensation (TDC) estimator and clustering analysis. The TDC estimator is a simultaneous operator to estimate binaural cues, which breaks the limitation of independent processors for binaural cues extraction. The multi-sources decision is realized by clustering analysis for the binaural cues of multiple signal frames. In experiments, we demonstrate that the localization performance is improved compared to the methods that assume the number of spatial stationary sources to be known. Results with both simulated and recorded impulse responses show that robust performance can be achieved with limited prior training, and our method is also adaptive to different sound activities.
Hong Liu 0008, Mengdi Yue, Jie Zhang 0042
IROS1
2016 A new descriptor of gradients Self-Similarity for smile detection in unconstrained scenarios
Yuan Gao 0008, Hong Liu 0008, Can Wang 0006
Neurocomputing2
2016 Depth Context: a new descriptor for human activity recognition by using sole depth sequences
Mengyuan Liu 0001, Hong Liu 0008
Neurocomputing2
2016 Human activity prediction by mapping grouplets to recurrent Self-Organizing Map
Qianru Sun, Hong Liu 0008, Tianwei Zhang 0002
Neurocomputing2
2016 A novel hierarchical Bag-of-Words model for compact action representation
Qianru Sun, Hong Liu 0008, Liqian Ma, Tianwei Zhang 0002
Neurocomputing2
2016 Violence detection using Oriented VIolent Flows
Yuan Gao 0008, Hong Liu 0008, Xiaohu Sun, Can Wang 0006
Image Vis. Comput.2
2016 A Novel Lip Descriptor for Audio-Visual Keyword Spotting Based on Adaptive Decision Fusion
abstract
Keyword spotting remains a challenge when applied to real-world environments with dramatically changing noise. In recent studies, audio-visual integration methods have demonstrated superiorities since visual speech is not influenced by acoustic noise. However, for visual speech recognition, individual utterance mannerisms can lead to confusion and false recognition. To solve this problem, a novel lip descriptor is presented involving both geometry-based and appearance-based features in this paper. Specifically, a set of geometry-based features is proposed based on an advanced facial landmark localization method. In order to obtain robust and discriminative representation, a spatiotemporal lip feature is put forward concerning similarities among textons and mapping the feature to intra-class subspace. Moreover, a parallel two-step keyword spotting strategy based on decision fusion is proposed in order to make the best use of audio-visual speech and adapt to diverse noise conditions. Weights generated using a neural network combine acoustic and visual contributions. Experimental results on the OuluVS dataset and PKU-AV dataset demonstrate that the proposed lip descriptor shows competitive performance compared to the state of the art. Additionally, the proposed audio-visual keyword spotting (AV-KWS) method based on decision-level fusion significantly improves the noise robustness and attains better performance than feature-level fusion, which is also capable of adapting to various noisy conditions.
Hong Liu 0008, Xiaofei Li 0001, Ting Fan, Xuewu Zhang 0003
IEEE Trans. Multim.2
2015 Body-structure based feature representation for person re-identification
abstract
Person re-identification is valuable for intelligent video surveillance and has drawn wide attention. Although person re-identification research is making progress, it still faces some challenges such as varying poses, illumination and viewpoints. As a major aspect of person re-identification, feature representation has been widely researched. Low-level descriptors are generally used in existing works, which do not take full advantage of body structure information and result in low discrimination. In this paper, body-structure based mid-level feature representation is proposed, which introduces body structure pyramid for codebook learning and feature pooling. Additionally, low computational LLC is used to encode mid-level features. Experimental results on two challenging datasets VIPeR and CUHK01 have demonstrated that our approach outperforms the state-of-the-art methods.
Hong Liu 0008, Liqian Ma, Can Wang 0006
ICASSP1
2015 Online person orientation estimation based on classifier update
abstract
Person orientation estimation is valuable for intelligent video surveillance. Although much progress has been made in recent years, it still faces challenges such as varying poses, illuminations and viewpoints. Most existing approaches merely use appearance information or combine it with motion information. Appearance-based classifiers are trained offline without updating in real time, which can not adapt to unknown scenes. To fix it, a novel orientation estimation approach based on online appearance-based classifier update is proposed. Reliable motion direction is determined acting as pre-estimated person orientation to update the appearance-based classifier. Moreover, a novel criterion based on motion reliability is proposed to determine the motion direction. Experimental results show that the proposed approach achieves more competitive performances especially for unknown scenes.
Hong Liu 0008, Liqian Ma
ICIP1
2015 SDM-BSM: A fusing depth scheme for human action recognition
abstract
Depth map has shown promising capability in human action recognition, however it always be auxiliary of RGB features in previous work. As to sufficiently exploring depth map, we propose an innovative descriptor for human action recognition using solo depth data. First, Salient Depth Map (SDM) is calculated between two consecutive depth frames, which is superior for action description as it is located on salient moving objects. Moreover, Binary Shape Map (BSM) is proposed to depict the silhouettes induced by the lateral component of the scene action parallel to the image plane. Then, for implementation, a new framework as Bag-of-Map-Words is employed after concatenating SDM and BSM feature vectors. Experiments on NHA database demonstrate the superiority and high efficiency of the proposed method. We also give detailed comparisons with other features and analysis for parameters as a guidance of further applications.
Hong Liu 0008, Mengyuan Liu 0001, Hao Tang 0005
ICIP1
2015 Regression based landmark estimation and multi-feature fusion for visual speech recognition
abstract
Visual speech recognition also known as lipreading can improve robustness of automatic acoustic speech recognition especially under noisy environments. However, it remains a challenging topic considering the variety of speaking characteristics and confusion between visual speech features. In this paper, we propose an automatic lipreading method by using a new lip tracking method and multiple visual information fusion to tackle the problem. First, a method of face landmark estimation based on regression is employed for lip detection, based on which a geometric-based shape invariant feature (SIF) is put forward. Moreover, it can also be applied to the removal of the non-speaking utterance. Then the motion interchange patterns and spatial-temporal descriptors are also adopted to describe the lip information, where the Bayes combination strategy is applied. The proposed method is explored on three benchmark data sets: Avletters2, OuluVS and PKUVS. Experimental results demonstrate promising results and show effectiveness of the proposed approach.
Hong Liu 0008, Xuewu Zhang 0003
ICIP1
2015 Two-level multi-task metric learning with application to multi-classification
abstract
Many metric learning approaches neglect that the real world multi-class problems share strong visual similarities, which can be exploited by learning discriminative models. In this paper, a Two-level Multi-task Metric Learning (TMTL) method is presented to learn a distance measure from equivalence constraints. Multiple features are adopted to represent the image information and learn the distance matrices in the first level. Then the task-specific learning paradigm and multi-task voting mechanism make full use of pairwise equivalence labels, which induces knowledge from anonymous pairs to multi-classification. Experiments are conducted on two challenging benchmarks PubFig and OuluVS for face identification and lipreading respectively. The results demonstrate that our method outperforms the recent multi-task learning approaches and multi-class support vector machine.
Hong Liu 0008, Xuewu Zhang 0003
ICIP1
2015 Binaural sound source localization based on generalized parametric model and two-layer matching strategy in complex environments
abstract
Binaural sound source localization is an important technique involving Human-Robot Interaction (HRI), video conference, speech enhancement, etc. In many real application scenarios, especially for closed environments, the affect of reverberation and noise would degrade the precision of position estimations. Therefore, a new binaural sound source localization method based on generalized parametric model and two-layer matching strategy is proposed in this paper for complex environments. Firstly, cepstral prefiltering is utilized for dereverberation of binaural signals. Then, two binaural cues computed from a dual-channel frequency representation, are combined to estimate the azimuths of sources. Additionally, the generalized parametric model is presented to describe the relationship between the azimuth and binaural cues through finding the optimal scaling factors from training data. At last, a two-layer matching strategy based on Bayesian rule is used to make the final decision, which can effectively decrease the computation complexity. Experiments have validated the proposed approach and show that it achieves favorably better results compared with several available methods without extra spacial burden.
Hong Liu 0008, Jie Zhang 0042
ICRA1
2015 A predictive model for narrow passage path planner by using Support Vector Machine in changing environments
abstract
Narrow passages in changing environments create huge difficulties, since locations and shapes of narrow passages in Configuration Space(C-space) change frequently. It is very important for a planner to identify narrow passages in real time and boost valid points within them effectively. A novel narrow passage predictive model for designing a path planner in changing environments is proposed in this paper. Firstly, an Expanded Dynamic Bridge Builder is presented to identify narrow passages rapidly with validity-toggle sampling points in C-space. Secondly, the predictive model is adopted to sample possibly free points within these narrow passages without invoking any collision detection in order to avoid intense computational complexity. The predictive model is obtained by the famous classification method of Support Vector Machine (SVM). A new feature, which includes a group of points' distance and validity information, is proposed in SVM training process to capture approximate structure of local narrow passages. Therefore, the predictive model can excavate the hidden similar structure of local narrow passages. Experiments carried out with two 6-DOFs manipulators show that our approach gain higher success rate of planning and time efficiency than other related methods.
Hong Liu 0008, Fang Xiao, Can Wang 0006
ICRA1
2015 Direction of arrival estimation based on reverberation weighting and noise error estimator
Jie Zhang 0042, Hong Liu 0008
INTERSPEECH3
2015 Gender Classification Using Pyramid Segmentation for Unconstrained Back-facing Video Sequences
abstract
This paper presents a pioneering study on gender classification from unconstrained back-facing video sequences in natural scenes. In many cases, classifying gender simply via faces or other biometric cues may fail when the video only contains back-facing people. To address this problem, we propose a novel approach to classify the gender according to back-facing video sequences. For this task, a novel Pyramid Segmentation approach is proposed to divide video sequence into a suite of equal time-length sleeves with different scales. Moreover, a heuristic approach is used to compute weights for different features from each sleeve. Finally, a framework of gender classification based on video sequences is presented. To validate our approach, we introduce a new dataset, called BackFacing dataset, featured by 720 annotated back-facing human video sequences. To our knowledge, this is the first dataset only containing back-facing video shots. Experiments demonstrate that the proposed approach achieves competitive results on VidTIMIT, Cohn-Kanade, CASIA Gait and BackFacing datasets.
Hao Tang 0005, Hong Liu 0008, Wei Xiao 0002
ACM Multimedia2
2015 Noise-free representation based classification and face recognition experiments
Yong Xu 0001, Xiaozhao Fang, Jane You, Yan Chen 0018, Hong Liu 0008
Neurocomputing5
2015 Exploring Spatial Correlation for Visual Object Retrieval
abstract
Bag-of-visual-words (BOVW)-based image representation has received intense attention in recent years and has improved content-based image retrieval (CBIR) significantly. BOVW does not consider the spatial correlation between visual words in natural images and thus biases the generated visual words toward noise when the corresponding visual features are not stable. This article outlines the construction of a visual word co-occurrence matrix by exploring visual word co-occurrence extracted from small affine-invariant regions in a large collection of natural images. Based on this co-occurrence matrix, we first present a novel high-order predictor to accelerate the generation of spatially correlated visual words and a penalty tree (PTree) to continue generating the words after the prediction. Subsequently, we propose two methods of co-occurrence weighting similarity measure for image ranking: Co-Cosine and Co-TFIDF. These two new schemes down-weight the contributions of the words that are less discriminative because of frequent co-occurrences with other words. We conduct experiments on Oxford and Paris Building datasets, in which the ImageNet dataset is used to implement a large-scale evaluation. Cross-dataset evaluations between the Oxford and Paris datasets and Oxford and Holidays datasets are also provided. Thorough experimental results suggest that our method outperforms the state of the art without adding much additional cost to the BOVW model.
Miaojing Shi, Xinghai Sun, Dacheng Tao, Chao Xu 0006, George Baciu, Hong Liu 0008
ACM Trans. Intell. Syst. Technol.6
2014 Learning directional co-occurrence for human action classification
abstract
Spatio-temporal interest point (STIP) based methods have shown promising results for human action classification. However, state-of-art works typically utilize bag-of-visual words (BoVW), which focuses on the statistical distribution of features but ignores their inherent structural relationships. To solve this problem, a descriptor, namely directional pair-wise feature (DPF), is proposed to encode the mutual direction information between pairwise words, aiming at adding more spatial discriminant to BoVW. Firstly, STIP features are extracted and classified into a set of labeled words. Then in each frame, the DPF is constructed for every pair of words with different labels, according to their assigned directional vector. Finally, DPFs are quantized to be a probability histogram as a representation of human action. The proposed method is evaluated on two challenging datasets, Rochester and UT-interaction, and the results based on chi-squared kernel SVM classifiers confirm that our method can classify human actions with high accuracies.
Hong Liu 0008, Qianru Sun
ICASSP1
2014 A binaural sound source localization model based on time-delay compensation and interaural coherence
abstract
Binaural sound source localization is an important technique involving speech capture and enhancement. However, the simple array structure makes it hard to localize sources in complex noisy conditions. This paper presents a novel algorithm based on time-delay compensation (TDC) and inter-aural coherence for binaural sound localization. Firstly, the TDC of binaural signals is used to estimate interaural time-delay (ITD) and interaural intensity difference (IID) instead of generalized cross correlation and logarithmic energy ratio. Then the interaural coherence is utilized to select reliable frames and reduce the variance of ITDs. Finally, a hierarchical framework, which successfully reduces computation complexity, is applied to make a decision of location based on Bayesian rule. Our innovation lies in that both ITD and IID are foremost yielded by TDC. Compared with other popular algorithms, experiments show that the most extrusive superiority of this method is complexity for both time and storage.
Hong Liu 0008, Jie Zhang 0042
ICASSP1
2014 Spontaneous versus posed smile recognition using discriminative local spatial-temporal descriptors
abstract
Automatic recognition of spontaneous versus posed (SVP) facial expressions has received widespread attention in recent years for its potential applications in friendly human machine interface. Most existing works of SVP facial expression recognition extract geometry-based features which heavily rely on accurate detection and tracking of facial feature points. In this paper, a novel approach is proposed to distinguish between spontaneous and posed smiles using discriminative completed LBP from three orthogonal planes, which is an appearance-based local spatial-temporal descriptor. The descriptor devotes to extracting most robust and discriminative patterns of interest. In addition, flexible facial subregion cropping, a spatial division method, is proposed taking into account different facial organ size of different people and filtering of redundant information. Besides, in the temporal domain, a new division method is also applied, which divides the smile process according to smile dynamics. Experiments on three benchmark databases and comparisons to the state-of-the-art methods validate the advantages of our approach, obtaining an accuracy rate of 91.40%.
Hong Liu 0008, Xuewu Zhang 0003
ICASSP2
2014 Smile detection in unconstrained scenarios using self-similarity of gradients features
abstract
Smile detection in unconstrained scenarios is a hot research topic with many real-world applications. This paper presents a new approach to practical smile detection and the primary contributions are three-fold. (1) In the image registration procedure, an eyes-mouth alignment strategy is found to be more efficient than popular eyes alignment. (2) In the feature extraction procedure, a novel feature descriptor, Self-Similarity of Gradients (GSS), is proposed and achieved good performance in comparison with baseline approaches. (3) Feature combination and multi-classifier combination strategies are adopted in experiments and excellent results are obtained. Experimental results show that the combined features (HOG+GSS) using AdaBoost+SVM achieve improved performance over state-of-the-art in the GENKI4K benchmark.
Hong Liu 0008, Yuan Gao 0008
ICIP1
2014 Gender identification in unconstrained scenarios using Self-Similarity of Gradients features
abstract
Gender identification has been a hot research topic with wide application requirements from social life. In general, effective feature representation is the key to solving this problem. In this paper, a new feature named Self-Similarity of Gradients (GSS) is proposed, which captures pairwise statistics of localized gradient distributions. There are three contributions made by us to practical gender identification. First, GSS features are proposed for gender identification in the wild, which achieve good performance compared with baseline approaches. Second, we originally utilize 31-dimensional HOG for practical gender identification and its excellent results demonstrates that HOG with both contrast sensitive and insensitive information is a better fit for this topic than that with only contrast insensitive information. Last, feature combination and multi-classifier combination strategies are adopted and the best gender identification performance is achieved. Experimental results show that the combination of GSS, HOG and LBP using a linear SVM outperforms state-of-the-art on the LFW database, which meets the “wild” condition.
Hong Liu 0008, Yuan Gao 0008, Can Wang 0006
ICIP1
2014 Action classification by exploring directional co-occurrence of weighted stips
abstract
Human action recognition is challenging mainly due to intro-variety, inter-ambiguity and clutter backgrounds in real videos. Bag-of-visual words model utilizes spatio-temporal interest points(STIPs), and represents action by the distribution of points which ignores visual context among points. To add more contextual information, we propose a method by encoding spatio-temporal distribution of weighted pairwise points. First, STIPs are extracted from an action sequence and clustered into visual words. Then, each word is weighted in both temporal and spatial domains to capture the relationships with other words. Finally, the directional relationships between co-occurrence pairwise words are used to encode visual contexts. We report state-of-the-art results on Rochester and UT-Interaction datasets to validate that our method can classify human actions with high accuracies.
Hong Liu 0008, Qianru Sun
ICIP2
2014 Audio-visual Keyword Spotting for Mandarin Based on Discriminative Local Spatial-Temporal Descriptors
abstract
Although keyword spotting (KWS) technologies have been successfully applied to some applications, most KWS systems have a common problem of noise-robustness when applied to real-world environments. Audio-visual keyword spotting (AVKWS) using both acoustic and visual information is a solution to complementarily solve the problem. Most existing audio-visual speech recognition (AVSR) systems extract geometric features as visual features, which heavily rely on accurate and reliable detection and tracking of facial feature points. To avoid this defect of geometric features, an appearance-based discriminative local spatial-temporal descriptor (disCLBP-TOP) is proposed in this paper, which devotes to extracting robust and discriminative patterns of interest. Besides, a parallel two-step recognition based on both acoustic and visual keyword searching and re-scoring is conducted, which complementarily makes the best of two modalities under different noisy conditions. Adaptive weights for decision fusion are generated using a sigmoid function based on reliabilities of the two modalities, capable of adapting to various noisy conditions. Experiments show that our proposed parallel AVKWS strategy based on decision fusion significantly improves the noise robustness and attains better performance than feature fusion based audio-visual spotter. Additionally, disCLBP-TOP shows more competitive performance than CLBP-TOP.
Hong Liu 0008, Ting Fan
ICPR1
2014 Audio-visual keyword spotting based on adaptive decision fusion under noisy conditions for human-robot interaction
abstract
Keyword spotting (KWS) deals with the identification of keywords in unconstrained speech, which is a natural, straightforward and friendly way for human-robot interaction (HRI). Most keyword spotters have the common problem of noise-robustness when applied to real-world environment with dramatically changing noises. Since visual information won't be affected by the acoustic noise, it can be utilized to complementarily improve the noise-robustness. In this paper, a novel audio-visual keyword spotting approach based on adaptive decision fusion under noisy conditions is proposed. In order to accurately represent the appearance and movement of mouth region, an improved local binary pattern from three orthogonal planes (ILBP-TOP) is proposed. Besides, a parallel two-step recognition based on acoustic and visual keyword candidates is conducted and generates corresponding acoustic and visual scores for each keyword candidate. Optimal weights for combining acoustic and visual contributions under diverse noise conditions are generated using a neural network based on reliabilities of the two modalities. Experiments show that our proposed audio-visual keyword spotting based on decision fusion significantly improves the noise robustness and attains better performance than feature fusion based audiovisual spotter. Additionally, ILBP-TOP shows more competitive performance than LBP-TOP.
Hong Liu 0008, Ting Fan
ICRA1
2014 A new hierarchical binaural sound source localization method based on Interaural Matching Filter
abstract
Binaural sound source localization is an important technique in friendly Human-Robot Interaction (HRI) for its easy-implementation with only two microphones. This paper develops a robust method based on a hierarchical probabilistic model. Reliable frequency sub-bands are used to PHAT — ργ method in the first layer to obtain a priori crude Interaural Time-Delay (ITD) and the probabilistic distribution of candidate azimuths. The second layer utilizes Interaural Intensity Difference (IID) to reduce matching time and refine candidate azimuths as well as elevations. A novel feature named Interaural Matching Filter (IMF), which can eliminate the difference between ITDs and IIDs, is proposed in the third layer. The probability of sound source location is acquired by computing the similarity between the IMF of received binaural signal and the IMFs in templates. Finally, combined with the former ITDs and IIDs, the similarity matrix is used to make a decision of sound source location based on a Bayes rule. Our innovation lies in adding selecting reliable frequency components into time-delay estimation and foremost taking IMF as a feature of sound source. Compared with several state-of-the-art algorithms, experimental results show our approach has a better performance even in noisy environments without increasing storages, and also has less time complexity.
Hong Liu 0008, Jie Zhang 0042, Zhuo Fu
ICRA1
2014 Locality and similarity preserving embedding for feature selection
Xiaozhao Fang, Yong Xu 0001, Xuelong Li 0001, Zizhu Fan, Hong Liu 0008, Yan Chen 0018
Neurocomputing5
2014 Modified minimum squared error algorithm for robust classification and face recognition experiments
Yong Xu 0001, Xiaozhao Fang, Qi Zhu 0001, Yan Chen 0018, Jane You, Hong Liu 0008
Neurocomputing6
2014 Inertial measurement unit-camera calibration based on incomplete inertial sensor information
abstract
This paper is concerned with the problem of estimating the relative orientation between an inertial measurement unit (IMU) and a camera. Unlike most existing IMU-camera calibrations, the main challenge in this paper is that the information output from the IMU is incomplete. For example, only two tilt information can be read from the gravity sensor of a smart phone. Despite incomplete inertial information, there are strong restrictions between the IMU and camera coordinate systems. This paper addresses the incomplete information based IMU-camera calibration problem by exploiting the intrinsic restrictions among the coordinate transformations. First, the IMU transformation between two poses is formulated with the unknown IMU information. Then the defective IMU information is restored using the complementary visual information. Finally, the Levenberg-Marquardt (LM) algorithm is applied to estimate the optimal calibration result in noisy environments. Experiments on both synthetic and real data show the validity and robustness of our algorithm.
Hong Liu 0008, Zhaopeng Gu
J. Zhejiang Univ. Sci. C1
2014 Contact-free and pose-invariant hand-biometric-based personal identification system using RGB and depth data
abstract
Hand-biometric-based personal identification is considered to be an effective method for automatic recognition. However, existing systems require strict constraints during data acquisition, such as costly devices, specified postures, simple background, and stable illumination. In this paper, a contactless personal identification system is proposed based on matching hand geometry features and color features. An inexpensive Kinect sensor is used to acquire depth and color images of the hand. During image acquisition, no pegs or surfaces are used to constrain hand position or posture. We segment the hand from the background through depth images through a process which is insensitive to illumination and background. Then finger orientations and landmark points, like finger tips or finger valleys, are obtained by geodesic hand contour analysis. Geometric features are extracted from depth images and palmprint features from intensity images. In previous systems, hand features like finger length and width are normalized, which results in the loss of the original geometric features. In our system, we transform 2D image points into real world coordinates, so that the geometric features remain invariant to distance and perspective effects. Extensive experiments demonstrate that the proposed hand-biometric-based personal identification system is effective and robust in various practical situations.
Can Wang 0006, Hong Liu 0008
J. Zhejiang Univ. Sci. C2
2014 Scene-Adaptive Hierarchical Data Association for Multiple Objects Tracking
abstract
Obtaining reliable and discriminative target representation are two vital tasks for data association in multi-tracking. Pervious works always directly combine bunch of features for more discriminative target representation, but this is prone to error accumulation and unnecessary computational cost, which on the contrary may increase identity switches in data association. Moreover, reliability of a same feature in different scenes may vary a lot, especially for currently widespread network cameras, which have been settled in complex and various scenes, previous fixed feature selection scheme cannot meet general requirements. To address this problem, we propose a scene-adaptive hierarchical data association scheme, which adaptively selects features which have higher reliability on target representation in applied scene, and gradually combines features to the minimum requirements of discriminating ambiguous targets. Hierarchical feature space is constructed according to reliability of features in the multi-tracking system, and data association is conducted in different layers of the feature space adaptively. Our algorithm is validated on various challenging RGB-D and RGB datasets recorded in various indoor and outdoor scenes, for diversities of both features and scenes. Experimental results validate its effectiveness and efficiency.
Can Wang 0006, Hong Liu 0008, Yuan Gao 0008
IEEE Signal Process. Lett.2
2014 Depth Motion Detection - A Novel RS-Trigger Temporal Logic based Method
abstract
Recently, depth data is widely used in computer vision applications such as detection and tracking, which shows great promises in complicated environments due to its complementary natures to RGB data. However, previous works mostly use depth as an auxiliary cue of RGB data and overlook its inherent advantage on motion detection. Intrinsically different from RGB data, points in depth map essentially represents 3-D positions in the world, so depth video represents the variation of these “positions,” which is motion. Motivated by this, we proposed a novel motion detection scheme based on RS-Trigger temporal logic which best fits nature of depth data on motion detection. The proposed algorithm can fast detect motion regions in the scene without statistics of background and prior knowledge of objects to detect. In following refinement modules, a depth-invariant density-constant projection is proposed which contributes to a fast spatial clustering and accurate segmentation, for it transforms dense 3-D points cloud to depth-invariant 2-D map with density-constance, not only it overcomes depth-dependent sampling of depth sensor, but also overcomes the common ‘scale problem’ in 2-D image analysis, which makes it easy to set system parameters to de-noise and pop-out motion regions. Experimental results validate its effectiveness and efficiency.
Can Wang 0006, Hong Liu 0008, Liqian Ma
IEEE Signal Process. Lett.2
2014 Data Uncertainty in Face Recognition
abstract
The image of a face varies with the illumination, pose, and facial expression, thus we say that a single face image is of high uncertainty for representing the face. In this sense, a face image is just an observation and it should not be considered as the absolutely accurate representation of the face. As more face images from the same person provide more observations of the face, more face images may be useful for reducing the uncertainty of the representation of the face and improving the accuracy of face recognition. However, in a real world face recognition system, a subject usually has only a limited number of available face images and thus there is high uncertainty. In this paper, we attempt to improve the face recognition accuracy by reducing the uncertainty. First, we reduce the uncertainty of the face representation by synthesizing the virtual training samples. Then, we select useful training samples that are similar to the test sample from the set of all the original and synthesized virtual training samples. Moreover, we state a theorem that determines the upper bound of the number of useful training samples. Finally, we devise a representation approach based on the selected useful training samples to perform face recognition. Experimental results on five widely used face databases demonstrate that our proposed approach can not only obtain a high face recognition accuracy, but also has a lower computational complexity than the other state-of-the-art approaches.
Yong Xu 0001, Xiaozhao Fang, Xuelong Li 0001, Jane You, Hong Liu 0008, Shaohua Teng
IEEE Trans. Cybern.6
2013 Inferring Ongoing Human Activities Based on Recurrent Self-Organizing Map Trajectory
abstract
Automatically inferring ongoing activities is to enable the early recognition of unfinished activities, which is quite meaningful for applications, such as online human-machine interaction and security monitoring. Stateof-the-art methods use the spatio-temporal interest point (STIP) based features as the low-level video description to handle complex scenes [1, 2, 3]. While the existing problem is that typical bag-of-visual words (BoVW) focuses on feature distribution but ignores the inherent contexts in sequences, resulting in low discrimination when directly dealing with limited observations. To solve this problem, the Recurrent SelfOrganizing Map (RSOM) [4], which was designed to process sequential data, is novelly adopted in this paper for the high-level representation of ongoing activities. The innovation lies that observed features and their spatio-temporal contexts are encoded in a trajectory of the pre-trained RSOM units. Additionally, a combination of Dynamic Time Warping (DTW) distance and Edit distance, named DTW-E, is specially proposed to measure the structural dissimilarity between RSOM trajectories. RSOM Trajectory: Since the RSOM constitutes a direct extension of SOM, we start from SOM. SOM is to map the data from an input space VI onto a lower dimensional space VL (a map) in such way that the topological relationships in VI are preserved and the SOM units approximate closely the probability density function of VI . Suppose each unit i in SOM is associated with a weight vector wi = [wi1,wi2, ...,win] ∈ Rn with the same dimension as the input vector x = [x1,x2, ...,xn] ∈ Rn. Learning process that leads to self-organization on a map can be summarized as, (i) The feature vector x(t) is input, then its best matching unit (bmu) on the map is found by computing the minimum distance as:
Qianru Sun, Hong Liu 0008
BMVC2
2013 Salient-motion-heuristic scheme for fast 3D optical flow estimation using RGB-D data
abstract
Optical flow is widely used for describing motion cues in the scene, but limited by slow estimating speed and illumination sensitivity. To handle both problems, this paper focuses on improving speed and accuracy of optical flow using RGB-D data and enhancing its robustness on motion description via fusing depth flow which is obtained only using depth data. First, salient motion regions (SMRs) are detected between depth frames which have good character on motion description for they all locate on moving objects. Then, depth flow is calculated to describe 3D motion for each SMR and directs fast orientation region growing on depth map. Thus larger motion regions are grown, and region-based optical flow estimation is conducted on grown regions. Estimation error is reduced and noise is inhibited due to depth constraints. Finally, a fusion scheme is adopted which combines depth flow and optical flow for better 3D motion description in the scene. Experiments on a RGB-D video data sets recorded in various complex scenes demonstrate the improved speed and robustness of the proposed method.
Can Wang 0006, Hong Liu 0008
ICASSP2
2013 Robust hand tracking based on online learning and multi-cue flocks of features
abstract
Robust hand tracking is in increasing demand from areas such as natural Human Robot/Computer interaction (HRI/HCI) and surveillance systems, while it is still a great challenge due to human hand's drastic appearance change. In recent years, online learning techniques have shown great potential in learning appearance of objects and tackling occlusion. This paper extends an online learning framework called Tracking-Learning-Detection (TLD) to track human hand. The main extensions are: 1) original tracker is replaced by a hybrid multi-cue base tracker combining Median-Flow tracker and Flock-of-Features (FoF) tracker, 2) skin color cue is integrated into cascade detector and PN online learning for more efficiency. Extensive experiments show that the new framework works with more robustness compared with state-of-the-art hand trackers and also original TLD.
Hong Liu 0008
ICIP1
2013 Hierarchical data association and depth-invariant appearance model for indoor multiple objects tracking
abstract
Discriminative target representation is vital for data association in multi-tracking. In order to increase the discriminative power, pervious works always combine bunch of features for target representation. However, this is prone to error accumulation and unnecessary computational cost, which may increase identity switches in data association on the contrary. To address this problem, we propose a hierarchical data association scheme which gradually combines features to the minimum requirements of discriminating ambiguous targets. In addition, indoor multi-tracking is more challenging due to frequent occlusion, view-truncation, large scale and pose variation, which may bring considerable unreliability for target representation. To handle this a novel depth-invariant part-based appearance model using RGB-D data is proposed. The depth-invariant appearance have stable length metric proportional to the absolute length metric in the world coordinates, which increase its robustness to scale variation. The part-based nature makes it robust to partial occlusion and view-truncation. Our algorithm is validated on various challenging indoor environments and it demonstrates high processing speed up to 50 fps and competitive accuracy.
Hong Liu 0008, Can Wang 0006
ICIP1
2013 Learning spatio-temporal co-occurrence correlograms for efficient human action classification
abstract
Spatio-temporal interest point (STIP) based features show great promises in human action analysis with high efficiency and robustness. However, they typically focus on bag-of-visual words (BoVW), which omits any correlation among words and shows limited discrimination in real-world videos. In this paper, we propose a novel approach to add the spatio-temporal co-occurrence relationships of visual words to BoVW for a richer representation. Rather than assigning a particular scale on videos, we adopt the normalized google-like distance (NGLD) to measure the words' co-occurrence semantics, which grasps the videos' structure information in a statistical way. All pairwise distances in spatial and temporal domain compose the corresponding NGLD correlograms, then their united form is incorporated with BoVW by training a multi-channel kernel SVM classifier. Experiments on real-world datasets (KTH and UCF sports) validate the efficiency of our approach for the classification of human actions.
Qianru Sun, Hong Liu 0008
ICIP2
2013 Unusual events detection based on multi-dictionary sparse representation using kinect
abstract
Unusual events detection plays a crucial role in surveillance applications, which is becoming more and more urgent need for public security. However, illumination and scale changing, lacking of sufficient training data and subjective of abnormality definition are some of the severe difficulties, which are hard to deal with by widely used traditional cameras. In order to solve these problems, first, a novel feature is proposed in this paper, which is named random local feature (RLF) to describe the spatial-temporal information of depth image detected by the Kinect sensor. Then, we expand the sparse representation framework to a multi-dictionary sparse representation framework, based on the intuition that that anomaly of a same event may vary a lot in different regions in a scene. We split the depth video into several regions and use detected RLF features in each region to train dictionary by K-SVD algorithm, and use the OMP algorithm to sparse-represent each feature. Finally, an objective function is introduced to evaluate the anomaly of features in each region according to reconstruction errors. Unusual events are defined as those incidences that occur very rarely in the entire video sequence in our system, which is tested on real data and demonstrates promising results in unusual events detection.
Can Wang 0006, Hong Liu 0008
ICIP2
2013 Maximally stable curvature regions for 3D hand tracking
abstract
Fast and robust hand detections and tracking is in increasing demand from areas such as natural Human Robot interaction(HRI) and surveillance systems. Previous works always use skin color or contour model to detection hand. However, they always fail for hands always exhibits drastic appearance change due to illumination change, non-rigid nature and hands are hard to discriminate from clutter background. Actually, the hand region has a specific nature that its curvature is relatively higher than other body parts and keeps stable whatever its poses and locations are, but none of pervious works exploit this nature for hand detection. In this work, a novel algorithm MSCR (Maximally Stable Curvature Regions) based on curvature nature to detect hands. It does not require manually initialization in the first frame, the hands are located by MSCR and skin color detector in the global image. 3D optical flow integrated Kalman Filter works to estimate the next location for local detector. Extensive experiments demonstrate that robust 3D tracking of hand articulations can be achieved in real-time with accurate results.
Can Wang 0006, Hong Liu 0008
ICIP2
2013 A two-layer probabilistic model based on time-delay compensation for binaural sound localization
abstract
Interaural Intensity Difference (IID) and Interaural Time Difference (ITD) are two important cues for robot acoustic localization both in Artificial Intelligence (AI) and Human-Robot Interaction (HRI) areas. However, it is a challenge job to localize a sound source accurately and swiftly only by two acoustic sensors. In this paper, a time-delay compensation based two-layer probabilistic model is presented for binaural sound source localization. In the first layer, a weighting function of Generalized Cross Correlation (GCC) named PHAT-ργ is used in low-frequency to obtain the prior time-delay. And in this layer a crude estimate of azimuth can also be acquired. At the same time, the probability of all possible time-delay lags can be achieved from the training data. In the Second layer, a new improved algorithm of IID based on time-delay compensation(named IIDδτ) is introduced to refine the probability of the azimuth and the elevation. Lastly, localization result is obtained by Bayes-Rule method. Comparing with three state-of-art algorithms, experimental results show that the proposed method has higher accuracy and costs less time for sound source localization.
Hong Liu 0008, Zhuo Fu, Xiaofei Li 0001
ICRA1
2013 Using the original and 'symmetrical face' training samples to perform representation based two-step face recognition
Yong Xu 0001, Xingjie Zhu, Guanghai Liu 0001, Yuwu Lu, Hong Liu 0008
Pattern Recognit.6
2013 Coarse to fine K nearest neighbor classifier
Yong Xu 0001, Qi Zhu 0001, Zizhu Fan, Minna Qiu, Yan Chen 0018, Hong Liu 0008
Pattern Recognit. Lett.6
2013 A comprehensive study on learning to rank for content-based image retrieval
Yangxi Li, Bo Geng, Chao Xu 0006, Hong Liu 0008
Signal Process.5
2013 Sound Source Localization for HRI Using FOC-Based Time Difference Feature and Spatial Grid Matching
abstract
In human-robot interaction (HRI), speech sound source localization (SSL) is a convenient and efficient way to obtain the relative position between a speaker and a robot. However, implementing a SSL system based on TDOA method encounters many problems, such as noise of real environments, the solution of nonlinear equations, switch between far field and near field. In this paper, fourth-order cumulant spectrum is derived, based on which a time delay estimation (TDE) algorithm that is available for speech signal and immune to spatially correlated Gaussian noise is proposed. Furthermore, time difference feature of sound source and its spatial distribution are analyzed, and a spatial grid matching (SGM) algorithm is proposed for localization step, which handles some problems that geometric positioning method faces effectively. Valid feature detection algorithm and a decision tree method are also suggested to improve localization performance and reduce computational complexity. Experiments are carried out in real environments on a mobile robot platform, in which thousands of sets of speech data with noise collected by four microphones are tested in 3D space. The effectiveness of our TDE method and SGM algorithm is verified.
Xiaofei Li 0001, Hong Liu 0008
IEEE Trans. Cybern.2
2013 Multiview Vector-Valued Manifold Regularization for Multilabel Image Classification
abstract
In computer vision, image datasets used for classification are naturally associated with multiple labels and comprised of multiple views, because each image may contain several objects (e.g., pedestrian, bicycle, and tree) and is properly characterized by multiple visual features (e.g., color, texture, and shape). Currently, available tools ignore either the label relationship or the view complementarily. Motivated by the success of the vector-valued function that constructs matrix-valued kernels to explore the multilabel structure in the output space, we introduce multiview vector-valued manifold regularization (MV(3)MR) to integrate multiple features. MV(3)MR exploits the complementary property of different features and discovers the intrinsic local geometry of the compact support shared by different features under the theme of manifold regularization. We conduct extensive experiments on two challenging, but popular, datasets, PASCAL VOC' 07 and MIR Flickr, and validate the effectiveness of the proposed MV(3)MR for image classification.
Yong Luo 0002, Dacheng Tao, Chang Xu 0002, Chao Xu 0006, Hong Liu 0008, Yonggang Wen 0001
IEEE Trans. Neural Networks Learn. Syst.5
2012 Action Disambiguation Analysis Using Normalized Google-Like Distance Correlogram
Qianru Sun, Hong Liu 0008
ACCV (3)2
2012 Comparison of methods for smile deceit detection by training AU6 and AU12 simultaneously
abstract
Smiles play an important role in face to face interaction which is a human-specific direct and naturally preeminent way of communication. People smile out of various reasons. Some smiles are spontaneous while some others are just posed. It is hard for human to catch such subtle differences between the two different smiles. In this paper, different algorithm combinations are tried to realize the deceit detection of smiles by computer, that is to say, to recognize true (spontaneous) and fake (posed) smiles. The detection is based on AUs (facial action units). Moreover, AU6 and AU12 are dealt with together in each example, which is different from AU recognition. Images in our database are all frontal facial images with smiles of different types and levels, which are from subjects of different countries and ages with different colors. Experiments are implemented to find out which algorithm combination is the best one. Results show that the best accuracy of the tried combinations in detecting true and fake smiles is close to 86%, while human true-fake-smile recognition ability is much lower. Our work could be used as a tool for the analysis of smiles in psychological area and other regions.
Hong Liu 0008
ICIP1
2012 Time Delay Estimation for Speech Signal Based on FOC-Spectrum
abstract
Higher-order statistics can be used for time delay estimation (TDE) to suppress spatially correlated Gaussian noise, since the higher-order cumulant of Gaussian signal is always zero. However, third-order statistics is invalid for those signals with zero skewness, speech signal as a typical one. In this paper, the fourth-order cumulant (FOC) spectrum is derived, based on which a TDE algorithm that is valid for speech signal and immune to spatially correlated Gaussian noise is proposed. This method can estimate the time delay between two sensor signals or simultaneously estimate the time delays between one sensor signal and other three. In addition, just like generalized cross correlation method, this spectrum domain algorithm is more robust than time domain FOC-based TDE algorithm, especially for speech signal due to its periodicity. Experiments verify the effectiveness of this TDE method for speech signal with spatially correlated Gaussian noise.
Hong Liu 0008, Xiaofei Li 0001
INTERSPEECH1
2012 Hierarchical RRT for humanoid robot footstep planning with multiple constraints in complex environments
abstract
Humanoid robots have abilities of stepping over or onto obstacles, which is different from wheeled robots. However it may be difficult to apply the ordinary motion planning methods such as Rapidly-exploring Random Trees (RRT) to humanoid robots directly. Because these kinds of methods only consider to circumvent obstacles and ignore the constraint of balance. Aiming at dealing with these problems in one frame, a novel approach based on hierarchical RRT is used to plan the footstep for humanoid robots. It is designed according to three basic constraints: a transition model based gait generator, an inverted pendulum based balance controller and a collision detection based path planner. First, a set of layered transition model is utilized to revise the footstep according to the terrain condition, which is able to take full use of the motion ability as well as improve the efficiency. Then, a hierarchical strategy is exploited to select the feasible foot location to be added in the random tree based on the results of collision checking and balance control. Finally, a dynamic RRT method is introduced in our work to revise paths in changing environments. Different experiments are given to verify the feasibility and performance of the proposed approach in complicated environments with both dynamic and static obstacles.
Hong Liu 0008, Tianwei Zhang 0002
IROS1
2012 A "capacitor" bridge builder based safe path planner for difficult regions identification in changing environments
abstract
Finding paths in difficult regions of C-space, such as narrow passages and configuration obstacle boundaries, is a rather challenging problem for path planning in changing environments. When obstacles move in W-space, these regions in C-space will change their edge points from free to collision or on the contrary, for which a “Capacitor” Bridge Builder (CBB) is proposed in this paper to identify their changing characteristics. Specifically, a “Capacitor” bridge is built between positive and negative toggled points in C-Space, which looks like capacitors stuck between narrow passages or boundary regions. Through CBB, the back boundary of an obstacle, which is less likely to be occupied immediately, is marked as a temporary safe region. Furthermore, a Half Bridge Strategy (HBS) is novelly proposed to boost samples inside these regions. Eventually, highly safe paths are revealed by predicting moving directions of obstacles, then replanning times and total planning times will be decreased significantly. Effectiveness of the proposed method has been verified by experiments with two manipulators in difficult changing environments.
Hong Liu 0008, Tianwei Zhang 0002, Chuangqi Wang
IROS1
2011 Real-time human tracking based on switching linear dynamic system combined with adaptive Meanshift tracker
abstract
Real-time human tracking in complex environments usually presents many challenges, such as partial/complete occlusions caused by irregular motion and similar-color distractors. Switching Linear Dynamic System(SLDS) and Meanshift(MS) are two successful approaches, although both have inherent deficiencies, like accumulated prediction and correction errors in SLDS, uncoded attentional and spatial information in Meanshift, respectively. In this paper, a spatial-representive Meanshift and a joint attentional feature histogram from bottom-up and top-down attention models are used to make the tracker more adaptive. Then a tracking algorithm is proposed as adaptive Meanshift embedded SLDS, where four transition matrixes A(st) handle partial/complete occlusions with irregular motion under the framework, and adaptive Meanshift solves short occlusions of similar-color distractors in local search for higher accuracy. Experiments show that this method can work more robustly for partial/complete occlusions among multiple persons compared with adaptive Meanshift and Meanshift embedded SLDS.
Hong Liu 0008, Chao Xu 0006
ICIP2
2011 Sound source localization for mobile robot based on time difference feature and space grid matching
abstract
Auditory is a convenient and efficient way for Human-Robot Interaction, however implementing a sound source localization system based on TDOA method encounters many problems, such as noise of real environments, and resolution of nonlinear equations, switch between far field and near field and lack of microphones for geometric positioning localization method. In this paper, a new spectral weighting GCC-PHAT method is proposed to deal with noise. Furthermore, the time difference feature of sound source and its spatial distribution are analyzed. Based on prosperities of the distribution, a space grid matching (SGM) algorithm is proposed for localization step, which handles those problems that geometric positioning method faces effectively. Decision tree and valid feature detection algorithm are also proposed to reduce computational complexity and improve performance. Experiments are achieved in real environments on a mobile robot platform, in which 2016 sets of speech data are tested using four microphones in 3D space. More than 95% azimuth localization rate with error less than 5 degrees and approximate 90% horizontal distance localization rate are obtained.
Xiaofei Li 0001, Hong Liu 0008, Xuesong Yang
IROS2
2010 A dynamic subgoal path planner for unpredictable environments
abstract
Although lots of planning algorithms have focused on the planning of fixed manipulators and mobile robots in moderate dynamic environments, seldom planning algorithms can be employed to deal with mobile agents in the presence of large scenario scales and unpredictable changing obstacles. Path planning for mobile robots in unpredictable environments would be an extreme challenge since computational complexity increase dramatically with high dimensionality, unpredictability and large scales. In this paper, a novel and real-time approach is proposed to solve this problem by generating subgoals dynamically according to time and potential values. This dynamic subgoal based approach includes two procedures, the subgoal generator and the inter-subgoal or inner replanner. On the one hand, a set of high-level subgoals is generated dynamically by an improved single shot strategy that could tailor itself adaptively. On the other hand, a roadmap is built during the preprocessing phase by employing a localized Dynamic Roadmap Mapping (local-DRM) for inter-subgoal replanning. Finally these two procedures will collaborate according to the potential field criterion to ensure completeness. Our approach can not only generate paths rapidly enough to satisfy the requirements of an anytime planer but also work for large scenario scales. Experimental results on different kinds of mobile agents, in large scenario scales and in the presence of unpredictable changing obstacles show that our approach can find out a collision free path on an average of 0.11s for a single planning, indicating an anytime planner.
Hong Liu 0008, Weiwei Wan, Hongbin Zha
ICRA1
2010 Continuous sound source localization based on microphone array for mobile robots
abstract
It is a great challenge to perform a sound source localization system for mobile robots, because noise and reverberation in room pose a severe threat for continuous localization. This paper presents a novel approach named guided spectro-temporal (ST) position localization for mobile robots. Firstly, since generalized cross-correlation (GCC) function based on time delay of arrival (TDOA) can not get accurate peak, a new weighting function of GCC named PHAT-ργ is proposed to weaken the effect of noise while avoiding intense computational complexity. Secondly, a rough location of sound source is obtained by PHAT-ργ method and room reverberation is estimated using such location as priori knowledge. Thirdly, ST position weighting functions are used for each cell in voice segment and all correlation functions from all cells are integrated to obtain a more optimistical location of sound source. Also, this paper presents a fast, continuous localization method for mobile robots to determine the locations of a number of sources in real-time. Experiments are performed with four microphones on a mobile robot. 2736 sets of data are collected for testing and more than 2500 sets of data are used to obtain accurate results of localization. Even if the noise and reverberation are serious. The proportion data is 92% with angle error less than 15 degrees. What's more, it takes less than 0.4 seconds to locate the position of sound source for each data.
Hong Liu 0008, Miao Shen
IROS1
2010 Adaptive replanning in hard changing environments
abstract
Replanning is a powerful tool for high dimensional mobile agents in changing environments. However, most works employ replanning periodically. In order to fully exert the merits of this powerful tool, we should concentrate on the time interval employed for each replanning (that is “when to replan”) and carry out replanning adaptively. In this paper, an adaptive strategy is proposed to govern replanning in hard changing environments. The key point of this adaptive replanning strategy is to perform local environment accumulation by using grids method, which is a derivative of degenerated potential field. Since the accumulation is only performed locally in the regions between subgoals and only computed towards the changes of obstacles, it increases little computational complexity to parent anytime planners. Our adaptive replanning strategy works as a plug-in to state-of-the-art algorithms and can generate heuristics by using information from projected spaces to overcome high dimensionality. Experiments on different mobile agents in various hard changing environments (environments with crowded and unforseen obstacles) with IDRM-gRRT and IRRT-gRRT showed that the adaptive strategy can improve the performance and robustness of parent anytime planners significantly.
Hong Liu 0008, Weiwei Wan
IROS1
2010 Omnidirectional vision for mobile robot human body detection and localization
abstract
Human body detection and localization is an essential capability of an autonomous mobile robot which works in the human-robot interaction (HRI) environments. However, due to field of view (FOV) limitations, it is hard to detect all human bodies around a mobile robot by using a conventional camera, and distances between robots and human bodies are also difficult to estimate. In this paper, we propose a novel omnidirectional visual system to locate positions of human bodies for an autonomous mobile robot. Firstly, a handy fitting shape based method (FSM) is presented to remap a omnidirectional image to a bird's eye view image. A new bird's eye view image segmentation algorithm, which is inspired by image pyramids, is used to split obstacle objects and ground plane. Secondly, a shape-based human body detector is implemented in unwrapped omnidirectional images to locate regions of human bodies. These human body detection results are combined with bird's eye view image segmentation to distinguish human bodies from other obstacle objects. Experiments show that our system performs well in human-robot interaction environments.
Hong Liu 0008, Zhenhua Huo, Ge Yang 0006
SMC1
2010 A selection method of speech vocabulary for human-robot speech interaction
abstract
Speech is the most natural and efficient way for Human-Robot Interaction (HRI), although speech recognition systems face some challenges on a mobile robot platform due to the wide range of users and varied noisy environments. This paper proposes a selection method of speech vocabulary for HRI, which can choose the most robust sub-vocabulary from the predefined isolated word vocabulary. We define a new concept, called Word Robustness, to represent the robustness of a word to speaker-independent and noise. The algorithm for computing Word Robustness is given based on Hidden Markov Model (HMM), then the most robust sub-vocabulary can be selected based on it. For convenience, this method makes use of selecting vocabulary to avoid other procedures such as speaker-adaption. Experiments were achieved based on an isolated word recognition system using a speech database which includes more than ten thousands of speech signals recorded in quiet laboratory and noisy environment respectively. Several sub-sets of vocabulary were selected for robot control based on Word Robustness. The best speaker-independent word recognition rate is 95.19% in noisy environments. Experimental results demonstrate the effectiveness of Word Robustness and the selection method.
Hong Liu 0008, Xiaofei Li 0001
SMC1
2009 Path planning in changing environments by using optimal path segment search
abstract
This paper presents a novel planner for manipulators and robots in changing environments. When environments are complicated, it's always difficult to find a completely valid path solution, which is essential for many methods. However, our planner searches for several path segments to make robot move towards its goal as much as possible even though such a complete solution doesn't exist currently. In the learning phase, the planner begins by building a roadmap that captures the topological structure of the configuration space in a workspace without obstacles. In the query phase, the planner searches for a solution path in the roadmap with the A* algorithm and performs roadmap updating using the lazy evaluation idea concurrently with the solution search process. If a completely valid solution is found, it will be adopted immediately. Otherwise the planner will collect a set of maximum valid path segments and then select the optimal one for planning in the execution process. The searching and execution process will be repeatedly performed until a goal configuration is reached. In plentiful experiments, our planner shows promising performances.
Hong Liu 0008
IROS1
2009 Collaboration of spatial and feature attention for visual tracking
abstract
Although primates can facilely maintain long-duration tracking of an object without infection of occlusion or other near similar distracters, it remains a challenge for computer vision system. Studies in psychology suggest that the ability of primates to focus selective attention on the spatial properties of an object is necessary to observe object quickly and efficiently while focus selective attention on the feature properties of object is necessary to render it more prominent from the distracters. In this paper, we propose a novel spatial-feature attentional visual tracking (SFAVT) algorithm to encode both. In SFAVT, tracking is treated as an on-line binary classification problem where spatial attention is employed in early selective procedure to construct foreground/background appearance model by identifying image patches with good localization properties, and in late selective procedure to update models by maintaining image patches with good discrimitive motion properties. Meanwhile, feature attention works in mode seeking procedure to help select feature spaces that best separate a target from background. The on-line tuned adaptive appearance models by those selected feature spaces are used to train a classifier for target localization, then. Experiments under various real-world conditions show that this algorithm is able to track an object influenced by dramatic distracters while is of comparable time efficiency with meanshift.
Hong Liu 0008, Weiwei Wan
IROS1
2009 Visual Tracking Algorithm Based on CAMSHIFT and Multi-cue Fusion for Human Motion Analysis
abstract
It is still a challenging problem for tracking objects in complex visual situations, such as an object is occluded or the object's color features are very similar to its background. Therefore, a novel visual tracking algorithm is proposed for multiple cues fusion based on three common cues: color, target position prediction and motion continuity in this paper. Color feature is free of translation and rotation and robust to partial occlusions and pose variations. Features of target position prediction and motion continuity can handle the condition that the color difference between the foreground and the background is similar. Combining with CAMSHIFT (Continuously Adaptive Mean Shift) technique, experimental results show that the proposed visual tracking algorithm is more robust than traditional single cue and gets better tracking effect than CMST (Collaborative Mean Shift Tracking). Successful rates of the proposed algorithm are 70% to 100% in 4 different complex conditions.
Ge Yang 0006, Hong Liu 0008
SMC2
2009 An Effective Background Reconstruction Method for Complicated Traffic Crossroads
abstract
Effective background reconstruction is the key for real time traffic flow monitoring. High traffic density and complexity of background scene make reconstruction more difficult. Background estimation based on the median method is imprecise under a complex traffic flow condition. In this paper, a new background estimation method based on the similarity of background using parameters of gray mean and variance is proposed. Therefore, a two-dimensional clustering and merging mechanism is introduced. At last, accurate decision about the category of the background is made by analyzing the distribution characteristic of the frame numbers in one category. Our algorithm works on the difficult condition of traffic congestion with higher reliability. The proposed method can be used in background reconstruction of the crossroads based on video sequences.
Hong Liu 0008
SMC1
2009 Image Restoration of Warped Complex Chinese Documents Based on Text Boundary Lines
abstract
Distortion always appears in document images while scanning thick bound volumes. There are two kinds of distortion for the scanned grayscale images, shadow appears at the volumes' spine area, and warping of the words occurs in the shadow. In this paper, a novel text boundary lines based method for efficient restoration of warped scanning Chinese document images is presented. We first detect on which side of an image the shadow lays by row grayscale analysis method. Then the shadow is removed by a modified Niblack's algorithm. In order to detect the warped feature, a text boundary lines' detection method is proposed. Finally, an adjustment method based on the text boundary lines is carried to restore the warped words. Experiments on 400 various scanning Chinese document images are implemented. The improvement on average character recall is 11.92% to 14.89%. Experiments show that the proposed restoration method is efficient for Chinese documents with both text and non-text regions.
Hong Liu 0008, Runwei Ding
SMC1
2009 A Method to Restore Chinese Warped Document Images Based on Binding Characters and Building Curved Lines
abstract
With rapid development of information technology, more and more document images are made by scanners. But new problem comes out that many of document images from thick books are warped. It is quite inconvenient for further process on computer. This paper introduces an integrative algorithm on restoring Chinese document images, which is a new filed and few researchers have worked on this subject yet. The complicated structure of Chinese ¿block words¿ makes the problem more difficult. To solve this, a restoring method which is based on binding characters iteratively and building curved lines using parallel lines method is introduced. In the phase of fitting, SVR is adopted instead of other parameter methods. An idea of collaboration is also recommended to guarantee the quality of the final results. Correction rate of 94% for experiment of 300 document images proves this method works out very well.
Hong Liu 0008
SMC1
2009 A Modified Cross Power-Spectrum Phase Method Based on Microphone Array for Acoustic Source Localization
abstract
The cross power-spectrum phase (CSP) method plays an important role in time delay estimate due to its efficiency in acoustic source localization. Generally, impacted by the spatial noise and reverberation in a room, the positioning accuracy could be greatly enhanced. However, this method could hardly lead to results when the energy of signal is small. In this paper, a new method is proposed: covariance matrix that from the multi-channel signals were applied and the characteristic of the coherent function was employed. Experiments under various conditions were carried out to demonstrate the new method's novelty, including adding background music to simulate the real environments. The 240 groups' data we achieved reached localization accuracy of 97.5% in normal experiments. We also obtained 85.42% and 83.75% correct rate in low and strongly musical environments, respectively.
Hong Liu 0008, Miao Shen
SMC1
2009 Combining Color Histogram and Gradient Orientation Histogram for Vision Based Global Localization
abstract
Global image features and local image features are comprehensively used in mobile robot's localization. In this paper, we proposed a geometric approach based on the combination of global image features. Considering the deficiency of the Weighted Gradient Orientation Histograms (WGOHs) for similar structure environments, color histograms are integrated as one vector for the localization. Besides the improving of weight and division for WGOH and color histogram, another weight for different emphasis on color and gradient orientation is carried out. A normalizing process is performed to better integrate the two global features. This combining approach is tested by means of locations recognition. Experimental results show that the proposed combining approach is efficient for indoor environments.
Hong Liu 0008, Xiaojia Yu
SMC1
2009 Motion Planning for Human-Robot Interaction Based on Stereo Vision and SIFT
abstract
It is very important for a robot to obverse its environment in real-time and walk without collision in a crowd. This paper presents a motion planning method, based on visual feedback, for safe Human-Robot Interaction (HRI) in dynamic environments. Firstly, in order to improve accuracy of features marching, Scale Invariant Feature Transform (SIFT) is merged into binocular stereo vision, which is used to detect motion of people. Secondly, by improving Lazy PRM, a robot can find the shortest safe path and move to predetermined destination along the path. Experimental results show that position of people can be detected in real-time in environments with several people walking inside, and the accuracy can reach 96%. Therefore, a robot can arrive at the goal configuration node without collision with people much faster than Lazy PRM.
Hong Liu 0008
SMC1
2009 Robust human tracking based on multi-cue integration and mean-shift
Hong Liu 0008, Hongbin Zha, Yuexian Zou
Pattern Recognit. Lett.1
2008 Adaptive feature-spatial representation for Mean-shift tracker
abstract
Mean-shift tracker plays an important role in computer vision applications due to its efficiency in mode seeking. By encoding the spatial information appropriately, the robustness of tracking could be greatly enhanced. However, to account for the deformation and other sources of variation of the tracking object, the spatial configuration should not be fixed apriori and it is more suitable to be adapted online. To this end, this paper presents a novel method to formulate an adaptive feature-spatial representation (FSR) for mean-shift tracking. By encoding blocking features of the tracking object with a set of adaptively weighted and spatially distributed tunable kernels, the object variations, like deformations and partial occlusions, can be handled appropriately. Extensive experiments under various conditions clearly demonstrate the obvious advantage of our approach compared to the classical mean-shift trackers.
Hong Liu 0008, Hongbin Zha
ICIP2
2008 Predictive model for path planning by using k-near dynamic bridge builder and Inner Parzen Window
abstract
Robotic path planning in changing environments with difficult regions is an extremely challenge. Since the structure of configuration space (C-space) will change when obstacles move in workspace (W-space), the planner should have the capacity of building approximate structure of C-space, while avoiding intense computational complexity. Further, difficult regions will also change their positions, which requires the planner should be able to identify them fast and increase the free nodes inside them efficiently. This paper presents a novel approach for path planning in changing environments using predictive model, which is inspired by the idea of active learning. With the help of W-C nodes mapping, this predictive model is built to capture the approximate structure of C-space, while avoiding intense computational complexity. This model include two steps: K-near Dynamic Bridge Builder (K-near DBB) is proposed to identify difficult passages in the space first, and then Inner Parzen Window is adopted to sample points in these difficult regions without invoking any collision checker. Experiments are carried out with two 6-DOF manipulators, and our approach can find a path with high time efficiency and low error rate, even if the environment is complex.
Hong Liu 0008, Weiwei Wan, Hongbin Zha
IROS1
2008 Skew detection for complex document images using robust borderlines in both text and non-text regions
Hong Liu 0008, Hongbin Zha
Pattern Recognit. Lett.1
2007 Fuzzy Decision Method for Motion Deadlock Resolving in Robot Soccer Games
Hong Liu 0008, Hongbin Zha
ICIC (1)1
2007 Collaborative Mean Shift Tracking Based on Multi-Cue Integration and Auxiliary Objects
abstract
Colour-based mean shift is an effective and fast algorithm for tracking colour blobs. However, it is vulnerable to full occlusion and target out of range for a few frames. This paper proposes a tracking method based on multi-cue integration and auxiliary objects to deal with these problems. A colour-location-prediction integration mean shift method is proposed to track each auxiliary object. Motivated by the idea of tuning weight of each cue according to their performances, these three cues are integrated adaptively according to their quality functions. Moreover, auxiliary objects get effective relative information with targets automatically, and update the information ceaselessly. When the target disappears, auxiliary objects will export useful information to estimate the location of the target. Experiments show that this method can adapt the weight of multi-cue efficiently, reinitialize the targets after long time disappearance, and increase the robustness of tracking in various conditions.
Hong Liu 0008, Hongbin Zha
ICIP (3)1
2007 A dynamic bridge builder to identify difficult regions for path planning in changing environments
abstract
This paper presents an efficient path planner to identify difficult regions for path planning in changing environments, in which obstacles can move randomly. The difficult regions consist of narrow passages and the boundaries of obstacles in robot Configuration Space (C-space). These regions exert significantly negative influence on finding a valid path in static environments. The problem becomes more complicated in changing environments, because that the regions will change their positions when obstacles move. Besides, it is necessary to identify difficult regions in real time since obstacles may move frequently. To identify difficult regions fast when they change their positions, a dynamic bridge builder is proposed based on a W-C nodes mapping and a Bridge planner method. The W-C nodes mapping is used not only to conserve the validity of nodes in C-space, but also to provide the information about where a "bridge" should be built, i.e. the positions of narrow passages, and where the boundaries of obstacles are. Furthermore, a hierarchy sampling strategy is employed to boost the density of nodes in difficult regions efficiently. In the query phase, a Lazy-edges evaluation method is adopted to validate the edges in a found path. Simulated experiments for a dual-manipulator system show that our method is efficient for path planning in changing environments.
Hong Liu 0008, Xuezhi Deng, Hongbin Zha
IROS2
2007 Automatic seal image retrieval method by using shape features of Chinese characters
abstract
In many eastern countries, a large number of seal images need to be identified every day. The system compares an input seal image with its reference seal and validates the authenticity of it. However, the reference seal is usually found manually. As each reference seal has quite a long corresponding ID, inputting ID to get the needed seal one by one costs quite a lot of time and this manual stage has become the bottleneck of automatic seal identification. To make the seal verification system more automatic, a new retrieval method based on Chinese characters' shape features will be introduced in this paper. Firstly, the main characters region is obtained by transforming the round seal image to a rectangle one and choosing the main part for each seal. Secondly, every single character is segmented using correlation method, and the number of characters in a seal can be got. Thirdly, four horizontal and four vertical features are extracted for each character in a seal and an eigenvector called position code is defined. Finally, the retrieval strategy is mainly based on transform from position code to a weight to decide which two seal images have the most similarity. Experiments on a database of 1000 testing seal images provide the retrieval ability for the proposed approach. The correct rate goes to 95.3%.
Hong Liu 0008, Hongbin Zha
SMC1
2006 A Path Planner in Changing Environments by Using W-C Nodes Mapping Coupled with Lazy Edges Evaluation
abstract
This paper presents a path planner based on PRM framework for robots operating in changing environments, in which obstacles can move randomly and robots may change their original shapes, e.g. a robot manipulator grasps an object. W-C nodes mapping coupled with lazy edges evaluation is used to ensure a generated path containing only valid nodes and edges when constructed probabilistic roadmap becomes partially invalid in changing environments. Our method combines merits of DRM and Lazy PRM methods. W-C nodes mapping, which is constructed in pre-processing phase by mapping every basic cell in workspace to nodes of roadmap, is preserved to indicate invalid nodes of roadmap fast whenever obstacles move. W-C edges mapping, which is another mapping relationship of DRM method, is skipped since it is much more complicated and time-consuming for construction. The simplified mapping with smaller size can be recomputed or modified fast when robots change their original shapes. Instead, lazy edges evaluation is used to ensure all edges valid along a found path. Simulated experiments for a dual-manipulator system show that our method is efficient for path planning in changing environments
Hong Liu 0008, Xuezhi Deng, Hongbin Zha
IROS1
2006 Robust Mean Shift Tracking Based on Multi-Cue Integration
abstract
Color-based mean shift has been addressed as an effective and fast algorithm for tracking color blobs. This deterministic searching method suffers from low saturation color object, color clutter in backgrounds and complete occlusion for several frames. This paper proposes a direct motion-color integration method to solve the low saturation color problem and the color background clutter problem. Based on the direct cue integration, an occlusion handler that is able to deal with long term full occlusion is proposed to solve the complete occlusion problem as well. Moreover, motivated by the idea of tuning weight of each cue according to its performance, a method of adaptive multi-cue integration based mean shift is proposed. Weights of each cue are adjusted according to a quality function, which is used to evaluate the performance of each cue in the adaptive integration scheme. Extensive experiments show that this method can adapt the weight of individual cue efficiently, and increase the robustness of tracking in various conditions.
Hong Liu 0008, Hongbin Zha
SMC1
2005 Document Image Retrieval Based on Density Distribution Feature and Key Block Feature
abstract
Document image retrieval is an important part of many document image processing systems such as paperless office systems, digital libraries and so on. Its task is to help users find out the most similar document images from a document image database. For developing a system of document image retrieval among different resolutions, different formats document images with hybrid characters of multiple languages, a new retrieval method based on document image density distribution features and key block features is proposed in this paper. Firstly, the density distribution and key block features of a document image are defined and extracted based on documents' print-core. Secondly, the candidate document images are attained based on the density distribution features. Thirdly, to improve reliability of the retrieval results, a confirmation procedure using key block features is applied to those candidates. Experimental results on a large scale document image database, which contains 10385 document images, show that the proposed method is efficient and robust to retrieve different kinds of document images in real time.
Hong Liu 0008, Suoqian Feng, Hongbin Zha
ICDAR1
2005 Real-Time and Distributed AV Content Analysis System for Consumer Electronics Networks
abstract
The ever-increasing complexity of generic multimedia-content-analysis-based (MCA) solutions, their processing power demanding nature and the need to prototype and assess solutions in a fast and cost-saving manner motivated the development of the Cassandra framework. The combination of state-of-the-art network and grid-computing solutions and recently standardized interfaces facilitated the set-up of this framework, forming the basis for multiple cross-domain and cross-organizational collaborations. It enables distributed computing scenario simulations for e.g. distributed content analysis (DCA) across consumer electronics (CE) in-home networks, but also the rapid development and assessment of complex multi-MCA-algorithm-based applications and system solutions. Furthermore, the framework's modular nature-logical MCA units are wrapped into so-called service units (SU)-ease the split between system-architecture- and algorithmic-related work and additionally facilitate reusability, extensibility and upgrade ability of those SUs
Jan Nesvadba, Pedro Fonseca 0002, Alexander Sinitsyn, Fons de Lange, Martijn Thijssen, Patrick van Kaam, Hong Liu 0008, Rien van Leeuwen, Johan J. Lukkien, Andrei Korostelev, Jan Ypma, Bart Kroon, Hasan Celik, Alan Hanjalic, Suphi Umut Naci, Jenny Benois-Pineau, Peter H. N. de With, Jungong Han
ICME7
2005 A planning method for safe interaction between human arms and robot manipulators
abstract
This paper presents a planning method based on mapping moving obstacles into C-space for safe interaction between human arms and robot manipulators. In pre-processing phase, a hybrid distance metric is defined to select neighboring sampled nodes in C-space to construct a roadmap. Then, two kinds of mapping are constructed to determine invalid and dangerous edges in the roadmap for each basic cell decomposed in workspace. For updating the roadmap when an obstacle is moving, basic cells covering the obstacle's surfaces are mapped into the roadmap by using new positions of the surfaces points sampled on the obstacle. In query phase, in order to predict and avoid coming collisions and reach the goal efficiently, an interaction strategy with six kinds of planning actions of searching, updating, walking, waiting, dodging and pausing are designed. Simulated experiments show that the proposed method is efficient for safe interaction between two working robot manipulators and two randomly moving human arms.
Hong Liu 0008, Xuezhi Deng, Hongbin Zha
IROS1
2005 Modeling facial expression space for recognition
abstract
In this paper, we present a method of modeling facial expression space for facial expression recognition by fuzzy integral. In traditional expression recognition methods using shape features, there are problems in describing both the uncertainty in facial expression classification and the relationship between facial features and facial expressions. Using facial expression space model, those problems can be solved easily. Firstly, we use values of fuzzy integral in different facial expression spaces to describe the uncertainty of facial expression. Secondly, by the fuzzy measure automatically constructed in each facial expression space, we deal with different effects of facial features for facial expression classification. Experiments show this method has a good ability of describing the uncertainty of facial expression and acquires good results of classification.
Hong Liu 0008, Hongbin Zha
IROS2
2005 Omni-directional vision based human motion detection for autonomous mobile robots
abstract
This paper presents a novel human motion detection system using an omni-camera on a platform of autonomous mobile robot. Compared with other related works, the system using an omni-directional camera provides a larger FOV (field of view) for motion detection, and makes a better effect using temporal differencing method based on compensation for ego-motion. Furthermore, the method, respectively combined with color information and feature point's on human body, is employed to determine human motion contours, which consequently improves the performance of the temporal differencing method. Experimental results show that the system provides an efficient way to track human motion in an indoor environment.
Hong Liu 0008, Hongbin Zha
SMC1
2004 3D model based head pose tracking by using weighted depth and brightness constraints
abstract
This paper proposes a robust method of tracking human head poses from a sequence of monocular images. First we estimate the head pose parameters in the first frame by an affine correspondence based method developed in our lab. Then both the linear brightness and depth constraint equations derived from the small interframe rigid motion assumption are used to implement the fast tracking of the head poses. We also take advantage of geometry information of the features on the face surface to weight the brightness and depth constraint equations to get more accurate results. Finally, in order to diminish the effects of gradual illumination changes and occlusions, we estimate the reliability of the features frame by frame and dynamically update the reliable feature set. Experiments show the proposed method can robustly track the head poses especially for the types of motions which make obvious depth variation.
Guoyuan Liang, Hongbin Zha, Hong Liu 0008
ICIG3