Runwei Ding

dblp:11/8205 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0003-4987-0405ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Debiased Multiplex Tokenizer for Efficient Map-Free Visual Relocalization
abstract
Image-based feature representation plays a critical role in visual localization, enabling robots to estimate their position and orientation in GPS-denied environments. However, this task is often undermined by significant variations in camera viewpoints and scene appearances. Recently, map-free visual relocalization (MFVR) has emerged as a promising paradigm due to its compatibility with lightweight deployment and privacy isolation on mobile devices. In this paper, we propose the Debiased Multiplex Tokenizer (DeMT) as a novel method for versatile and efficient MFVR. Specifically, DeMT performs relative pose regression through an integrated framework built upon a pretrained vision Mamba encoder, comprising three key modules: First, Multiplex Interactive Tokenization yields robust image tokens with non-local affinities and cross-domain descriptions; Second, Debiased Anchor Registration facilitates anchor token matching through proximity graph retrieval and causal pointer attribution; Third, Geometry-Informed Pose Regression empowers multi-layer perceptrons with a gating mechanism and spectral normalization to support both pair-wise and multi-view modes. Extensive evaluations across nine public datasets demonstrate that DeMT substantially outperforms existing baselines and ablation variants in diverse indoor and outdoor environments.
Hong Liu 0008, Shengquan Li 0001, Peifeng Jiang, Runwei Ding
AAAI5
2026 Turbo principles meet compression: Rethinking nonlinear transformations in learned image compression
Chao Li 0071, Wen Tan 0001, Fanyang Meng, Runwei Ding, Ye Wang 0002, Wei Liu 0065, Yongsheng Liang 0001
J. Vis. Commun. Image Represent.4
2025 Progressive Diffusion-Based Low Rate Perceptual Image Compression with Discrete Gaussian Codebooks for Remote Sensing
Yangxuan Cheng, Fanyang Meng, Runwei Ding, Ye Wang 0002, Yongsheng Liang 0001
PRCV (9)4
2024 Cloth Interactive Transformer for Virtual Try-On
abstract
The 2D image-based virtual try-on has aroused increased interest from the multimedia and computer vision fields due to its enormous commercial value. Nevertheless, most existing image-based virtual try-on approaches directly combine the person-identity representation and the in-shop clothing items without taking their mutual correlations into consideration. Moreover, these methods are commonly established on pure convolutional neural networks (CNNs) architectures which are not simple to capture the long-range correlations among the input pixels. As a result, it generally results in inconsistent results. To alleviate these issues, in this article, we propose a novel two-stage cloth interactive transformer (CIT) method for the virtual try-on task. During the first stage, we design a CIT matching block, aiming at precisely capturing the long-range correlations between the cloth-agnostic person information and the in-shop cloth information. Consequently, it makes the warped in-shop clothing items look more natural in appearance. In the second stage, we put forth a CIT reasoning block for establishing global mutual interactive dependencies among person representation, the warped clothing item, and the corresponding warped cloth mask. The empirical results, based on mutual dependencies, demonstrate that the final try-on results are more realistic. Substantial empirical results on a public fashion dataset illustrate that the suggested CIT attains competitive virtual try-on performance.
Bin Ren 0005, Hao Tang 0005, Fanyang Meng, Runwei Ding, Philip Torr 0001, Nicu Sebe
ACM Trans. Multim. Comput. Commun. Appl.4
2023 HTNet: Human Topology aware network for 3d Human pose estimation
abstract
3D human pose estimation errors would propagate along the human body topology and accumulate at the end joints of limbs. Inspired by the backtracking mechanism in automatic control systems, we design an Intra-Part Constraint module that utilizes the parent nodes as the reference to build topological constraints for end joints at the part level. Further considering the hierarchy of the human topology, joint-level and body-level dependencies are captured via graph convolutional networks and self-attentions, respectively. Based on these designs, we propose a novel Human Topology aware Network (HTNet), which adopts a channel-split progressive strategy to sequentially learn the structural priors of the human topology from multiple semantic levels: joint, part, and body. Extensive experiments show that the proposed method improves the estimation accuracy by 18.7% on the end joints of limbs and achieves state-of-the-art results on Human3.6M and MPI-INF-3DHP datasets. Code is available at https://github.com/vefalun/HTNet.
Jialun Cai, Hong Liu 0008, Runwei Ding, Wenhao Li 0002, Jianbing Wu, Miaoju Ban
ICASSP3
2023 Interweaved Graph and Attention Network for 3D Human Pose Estimation
abstract
Despite substantial progress in 3D human pose estimation from a single-view image, prior works rarely explore global and local correlations, leading to insufficient learning of human skeleton representations. To address this issue, we propose a novel Interweaved Graph and Attention Network (IGANet) that allows bidirectional communications between graph convolutional networks (GCNs) and attentions. Specifically, we introduce an IGA module, where attentions are provided with local information from GCNs and GCNs are injected with global information from attentions. Additionally, we design a simple yet effective U-shaped multi-layer perceptron (uMLP), which can capture multi-granularity information for body joints. Extensive experiments on two popular benchmark datasets (i.e. Human3.6M and MPI-INF-3DHP) are conducted to evaluate our proposed method. The results show that IGANet achieves state-of-the-art performance on both datasets. Code is available at https://github.com/xiu-cs/IGANet.
Ti Wang, Hong Liu 0008, Runwei Ding, Wenhao Li 0002, Yingxuan You, Xia Li 0005
ICASSP3
2023 Gator: Graph-Aware Transformer with Motion-Disentangled Regression for Human Mesh Recovery from a 2D Pose
abstract
3D human mesh recovery from a 2D pose plays an important role in various applications. However, it is hard for existing methods to simultaneously capture the multiple relations during the evolution from skeleton to mesh, including joint-joint, joint-vertex and vertex-vertex relations, which often leads to implausible results. To address this issue, we propose a novel solution, called GATOR, that contains an encoder of Graph-Aware Transformer (GAT) and a decoder with Motion-Disentangled Regression (MDR) to explore these multiple relations. Specifically, GAT combines a GCN and a graph-aware self-attention in parallel to capture physical and hidden joint-joint relations. Furthermore, MDR models joint-vertex and vertex-vertex interactions to explore joint and vertex relations. Based on the clustering characteristics of vertex offset fields, MDR regresses the vertices by composing the predicted base motions. Extensive experiments show that GATOR achieves state-of-the-art performance on two challenging benchmarks. Code is available at https://github.com/kasvii/GATOR.
Yingxuan You, Hong Liu 0008, Xia Li 0005, Wenhao Li 0002, Ti Wang, Runwei Ding
ICASSP6
2023 Multi-Stream Facial Adaptive Network for Expression Recognition from a Single Image
abstract
Facial expression recognition from a single image has potential applications in fields including human-computer interaction and medical diagnosis. Most recent methods use deep neural networks to directly learn from a roughly cropped facial image which is usually detected from a whole image by face detection algorithms. We observe that unrelated surrounding regions in the rough facial image prevent deep neural networks from learning facial-related discriminate features. To solve this problem, we present a Facial Adaptive Network (FAN) which is able to adaptively select an interest region from the given facial image, thus suffering less from the effect of unrelated regions. Based on the selected interest region, we further apply the self-attention mechanism to learn discriminate facial features. Moreover, we introduce a multi-stream FAN (ms-FAN) that learns richer facial features from multiple interest regions that are selected from pose-augmented facial images. Extensive experiments on Oulu-CASIA, CK+, and RAF-DB datasets consistently verify the effect of our proposed MS-FAN by achieving comparable results with state-of-the-art methods. Our code is available at https://github.com/zhangbc12/DAtt-ViT.
Baichuan Zhang, Fanyang Meng, Runwei Ding, Mengyuan Liu 0001
ICASSP3
2023 Co-Evolution of Pose and Mesh for 3D Human Body Estimation from Video
abstract
Despite significant progress in single image-based 3D human mesh recovery, accurately and smoothly recovering 3D human motion from a video remains challenging. Existing video-based methods generally recover human mesh by estimating the complex pose and shape parameters from coupled image features, whose high complexity and low representation ability often result in inconsistent pose motion and limited shape patterns. To alleviate this issue, we introduce 3D pose as the intermediary and propose a Pose and Mesh Co-Evolution network (PMCE) that decouples this task into two parts: 1) video-based 3D human pose estimation and 2) mesh vertices regression from the estimated 3D pose and temporal image feature. Specifically, we propose a two-stream encoder that estimates mid-frame 3D pose and extracts a temporal image feature from the input image sequence. In addition, we design a co-evolution decoder that performs pose and mesh interactions with the image-guided Adaptive Layer Normalization (AdaLN) to make pose and mesh fit the human body shape. Extensive experiments demonstrate that the proposed PMCE outperforms previous state-of-the-art methods in terms of both per-frame accuracy and temporal consistency on three benchmark datasets: 3DPW, Human3.6M, and MPI-INF-3DHP. Our code is available at https://github.com/kasvii/PMCE.
Yingxuan You, Hong Liu 0008, Ti Wang, Wenhao Li 0002, Runwei Ding, Xia Li 0005
ICCV5
2023 Achieving domain generalization for underwater object detection by domain mixup and contrastive learning
Pinhao Song, Hong Liu 0008, Linhui Dai, Xiaochuan Zhang, Runwei Ding, Shengquan Li 0001
Neurocomputing6
2023 Weakly-Supervised 3D Human Pose Estimation With Cross-View U-Shaped Graph Convolutional Network
abstract
Although monocular 3D human pose estimation methods have made significant progress, it is far from being solved due to the inherent depth ambiguity. Instead, exploiting multi-view information is a practical way to achieve absolute 3D human pose estimation. In this paper, we propose a simple yet effective pipeline for weakly-supervised cross-view 3D human pose estimation. By only using two camera views, our method can achieve state-of-the-art performance in a weakly-supervised manner, requiring no 3D ground truth but only 2D annotations. Specifically, our method contains two steps: triangulation and refinement. First, given the 2D keypoints that can be obtained through any classic 2D detection methods, triangulation is performed across two views to lift the 2D keypoints into coarse 3D poses. Then, a novel cross-view U-shaped graph convolutional network (CV-UGCN), which can explore the spatial configurations and cross-view correlations, is designed to refine the coarse 3D poses. In particular, the refinement progress is achieved through weakly-supervised learning, in which geometric and structure-aware consistency checks are performed. We evaluate our method on the standard benchmark dataset, Human3.6M. The Mean Per Joint Position Error on the benchmark dataset is 27.4 mm, which outperforms existing state-of-the-art methods remarkably (27.4 mm vs 30.2 mm).
Guoliang Hua, Hong Liu 0008, Wenhao Li 0002, Runwei Ding, Xin Xu 0001
IEEE Trans. Multim.5
2023 Exploiting Temporal Contexts With Strided Transformer for 3D Human Pose Estimation
abstract
Despite the great progress in 3D human pose estimation from videos, it is still an open problem to take full advantage of a redundant 2D pose sequence to learn representative representations for generating one 3D pose. To this end, we propose an improved Transformer-based architecture, called Strided Transformer, which simply and effectively lifts a long sequence of 2D joint locations to a single 3D pose. Specifically, a Vanilla Transformer Encoder (VTE) is adopted to model long-range dependencies of 2D pose sequences. To reduce the redundancy of the sequence, fully-connected layers in the feed-forward network of VTE are replaced with strided convolutions to progressively shrink the sequence length and aggregate information from local contexts. The modified VTE is termed as Strided Transformer Encoder (STE), which is built upon the outputs of VTE. STE not only effectively aggregates long-range information to a single-vector representation in a hierarchical global and local fashion, but also significantly reduces the computation cost. Furthermore, a full-to-single supervision scheme is designed at both full sequence and single target frame scales applied to the outputs of VTE and STE, respectively. This scheme imposes extra temporal smoothness constraints in conjunction with the single target frame supervision and hence helps produce smoother and more accurate 3D poses. The proposed Strided Transformer is evaluated on two challenging benchmark datasets, Human3.6 M and HumanEva-I, and achieves state-of-the-art results with fewer parameters. Code and models are available athttps://github.com/Vegetebird/StridedTransformer-Pose3D.
Wenhao Li 0002, Hong Liu 0008, Runwei Ding, Mengyuan Liu 0001, Pichao Wang, Wenming Yang
IEEE Trans. Multim.3
2022 Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action Recognition
abstract
In recent years, self-supervised representation learning for skeleton-based action recognition has been developed with the advance of contrastive learning methods. The existing contrastive learning methods use normal augmentations to construct similar positive samples, which limits the ability to explore novel movement patterns. In this paper, to make better use of the movement patterns introduced by extreme augmentations, a Contrastive Learning framework utilizing Abundant Information Mining for self-supervised action Representation (AimCLR) is proposed. First, the extreme augmentations and the Energy-based Attention-guided Drop Module (EADM) are proposed to obtain diverse positive samples, which bring novel movement patterns to improve the universality of the learned representations. Second, since directly using extreme augmentations may not be able to boost the performance due to the drastic changes in original identity, the Dual Distributional Divergence Minimization Loss (D3M Loss) is proposed to minimize the distribution divergence in a more gentle way. Third, the Nearest Neighbors Mining (NNM) is proposed to further expand positive samples to make the abundant information mining process more reasonable. Exhaustive experiments on NTU RGB+D 60, PKU-MMD, NTU RGB+D 120 datasets have verified that our AimCLR can significantly perform favorably against state-of-the-art methods under a variety of evaluation protocols with observed higher quality action representations. Our code is available at https://github.com/Levigty/AimCLR.
Tianyu Guo 0001, Hong Liu 0008, Mengyuan Liu 0001, Runwei Ding
AAAI6
2022 PDD-Net: A Precise Defect Detection Network Based on Point Set Representation
abstract
Defect detection has been widely studied in computer vision and used in industrial production. However, most existing methods for defect detection mainly suffer three drawbacks: i) Low-contrast problem between defects and background. ii) Large scale changes in defects size. iii) Extreme imbalance problem between defects and background classes during training. To address these issues, we propose a novel anchor-free defect detection network named PDD-Net. Specifically, a global-context FPN (GC-FPN) is designed to capture long-range dependency between defects and background. Simultaneously, to enhance feature extraction of defects at different scales, a receptive field pyramid block (RFPB) is proposed to provide various receptive field sizes. Furthermore, an equipped adaptive positive and negative samples allocation (APNSA) mechanism is built with statistical characteristics of defects, thus can select training samples automatically. We conduct experiments on MPSD dataset, DAGM2007 dataset, and NEU-DET dataset. Extensive experimental results on the three challenging datasets show that our PDD-Net achieves superior detection accuracy over the state-of-the-art methods.
Miaoju Ban, Runwei Ding, Jian Zhang 0117, Tianyu Guo 0001
ICASSP2
2022 Adaptive Weighted Network With Edge Enhancement Module For Monocular Self-Supervised Depth Estimation
abstract
Monocular self-supervised depth estimation can be easily applied in many areas since only a single camera is required. However, current methods do not predict well in depth borders. Besides, factors such as occlusion and texture sparsity can lead to the failure of the photometric consistency, affecting the prediction performance. To overcome these deficiencies, an adaptive weighted monocular self-supervised depth estimation framework that exploits enhanced edge information and texture sparsity based adaptive weights is proposed. In particular, a module named edge enhancement module (EEM) is designed to be embedded into the current depth prediction network to extract edge details for clearer depth prediction in depth borders. Moreover, a texture sparsity based adaptive weighted (TSAW) loss is introduced to as-sign different weights according to texture sparsity, enabling a more targeted construction of geometric constraints. Experimental results on the KITTI dataset demonstrate that the proposed network outperforms state-of-the-art methods.
Hong Liu 0008, Guoliang Hua, Weibo Huang, Runwei Ding
ICASSP5
2022 FDSNeT: An Accurate Real-Time Surface Defect Segmentation Network
abstract
Surface defect detection is a common task for industrial quality control, which increasingly requires accuracy and real-time ability. However, the current segmentation networks are not effective in dealing with defect boundary details, local similarity of different defects and low contrast between defect and background. To this end, we propose a real-time surface defect segmentation network (FDSNet) based on two-branch architecture, in which two corresponding auxiliary tasks are introduced to encode more boundary details and semantic context. To handle the local similarity problem of different surface defects, we propose a Global Context Upsampling (GCU) module by capturing long-range context from multi-scales. Moreover, we present a representative Mobile phone screen Surface Defect (MSD) segmentation dataset to alleviate the lack of dataset in this field. Experiments on NEU-Seg, Magnetic-tile-defect-datasets and MSD dataset show that the proposed FDSNet achieves promising trade-off between accuracy and inference speed. The dataset and code are available at https://github.com/jianzhang96/fdsnet.
Jian Zhang 0117, Runwei Ding, Miaoju Ban, Tianyu Guo 0001
ICASSP2
2022 HMFCA-Net: Hierarchical multi-frequency based Channel attention net for mobile phone surface defect detection
Runwei Ding, Weibo Huang
Pattern Recognit. Lett.2
2020 Grouped Temporal Enhancement Module for Human Action Recognition
abstract
Temporal information is a significant cue for recognizing human actions from videos. Different from 2D CNN which can only capture spatial information in an efficient way, 3D CNN is good at capturing both spatial and temporal information at the expense of high computational cost. Beyond both methods, this paper presents a Grouped Temporal Enhancement (GTE) module which even outperforms 3D CNN, meanwhile only needs similar low computational cost as 2D CNN. The GTE module firstly decomposes an input video into spatial and temporal groups along channel dimension, and then uses a learnable temporal shift (LTS) operation for efficient temporal modeling. Finally, a 2D convolution filter is used to enhance the ability of LTS for spatial modeling. Extensive experiments on three benchmark datasets validate the effect of our method.
Hong Liu 0008, Bin Ren 0005, Mengyuan Liu 0001, Runwei Ding
ICIP4
2020 Towards Domain Generalization In Underwater Object Detection
abstract
A General Underwater Object Detector (GUOD) should perform well on most of underwater circumstances. However, with limited underwater dataset, conventional object detection methods suffer from domain shift severely. This paper aims to build a GUOD using small underwater dataset with limited types of water quality. First, we propose a data augmentation method Water Quality Transfer (WQT) to increase domain diversity of the original small dataset. Second, for mining the semantic information from data generated by WQT, Domain Generalization YOLO (DG-YOLO) is proposed, which consists of three parts: YOLOv3, Domain Invariant Module and Invariant Risk Minimization penalty. Finally, experiments on original and synthetic URPC2019 dataset prove that WQT combined with DG-YOLO achieves promising performance of domain generalization in underwater object detection. The source code can be found at https://github.com/mousecpn/DG-YOLO.
Hong Liu 0008, Pinhao Song, Runwei Ding
ICIP3
2020 EDD-Net: An Efficient Defect Detection Network
abstract
As the most commonly used communication tool, the mobile phone has become an indispensable part of our daily life. The surface of the mobile phone as the main window of human-phone interaction directly affects the user experience. It is necessary to detect surface defects on the production line in order to ensure the high quality of the mobile phone. However, the existing mobile phone surface defect detection is mainly done manually. Currently, there are few automatic defect detection methods to replace human eyes. How to quickly and accurately detect the surface defects of the mobile phone is an urgent problem to be solved. Hence, an efficient defect detection network (EDD-Net) is proposed. Firstly, EfficientNet is used as the backbone network. Then, according to the small-scale of mobile phone surface defects, a feature pyramid module named GCSA-BiFPN is proposed to obtain more discriminative features. Finally, the box/class prediction network is used to achieve effective defect detection. We also build a mobile phone surface oil stain defect (MPSOSD) dataset to alleviate the lack of dataset in this field. The performance on the relevant datasets shows that the proposed network is effective and has practical significance for industrial production.
Tianyu Guo 0001, Runwei Ding
ICPR3
2020 A Base-Derivative Framework for Cross-Modality RGB-Infrared Person Re-Identification
abstract
Cross-modality RGB-infrared (RGB-IR) person reidentification (Re-ID) is a challenging research topic due to the heterogeneity of RGB and infrared images. In this paper, we aim to find some auxiliary modalities, which are homologous with the visible or infrared modalities, to help reduce the modality discrepancy caused by heterogeneous images. Accordingly, a new base-derivative framework is proposed, where base refers to the original visible and infrared modalities, and derivative refers to the two auxiliary modalities that are derived from base. In the proposed framework, the double-modality cross-modal learning problem is reformulated as a four-modality one. After that, the images of all the base and derivative modalities are fed into the feature learning network. With the doubled input images, the learned person features become more discriminative. Furthermore, the proposed framework is optimized by the enhanced intra- and cross-modality constraints with the assistance of two derivative modalities. Experimental results on two publicly available datasets SYSU-MM and RegDB show that the proposed method outperforms the other state-of-the-art methods. For instance, we achieve a gain of over 13 % in terms of both Rank- and mAP on RegDB dataset.
Hong Liu 0008, Ziling Miao, Bing Yang 0004, Runwei Ding
ICPR4
2020 Mobile Phone Surface Defect Detection Based on Improved Faster R-CNN
abstract
Various surface defects will inevitably occur in the production process of mobile phones, which have a huge impact on the enterprise. Therefore, precise defect detection is of great significance in the production of mobile phones. However, traditional manual inspection and machine vision inspection have low efficiency and accuracy respectively which cannot meet the rapid production needs of modern enterprises. In this paper, we proposed a mobile phone surface defect (MPSD) detection model based on deep learning, which greatly reduces the requirement of a large dataset and improves detection performance. First, Boundary Equilibrium Generative Adversarial Networks (BEGAN) is used to generate and augment the defect data. Then, based on the Faster R-CNN model, Feature Pyramid Network (FPN) and ResNet 101 are combined as feature extraction network to get more small target defect features. Further, replacing the ROI pooling layer with an ROI Align layer reduces the quantization deviation during the pooling process. Finally, we train and evaluate our model on our own dataset. The experimental results indicate that compared with some traditional methods based on handcrafted feature extraction and the traditional Faster R-CNN, the improved Faster R-CNN achieves 99.43% mAP, which is more effective in the field of MPSD defect detection.
Can Zhang 0007, Runwei Ding
ICPR3
2019 End-To-End Visual Place Recognition Based on Deep Metric Learning and Self-Adaptively Enhanced Similarity Metric
abstract
Place recognition, which aims at recognizing the previously visited places, is the key component of loop closure in most visual simultaneous localization and mapping systems. Despite significant progress, challenges still remain especially in the longtime mapping and localization as the appearance of an environment may change greatly over time. In this paper, deep metric learning which jointly optimizes feature extraction and similarity metric is utilized to train an end-to-end network specifically for place recognition task to handle the appearance changing over time. A self-adaptively enhanced similarity metric is designed to strength the discrimination ability and calculate the similarity between descriptors of image pairs which are extracted from a convolutional neural network. Experiments on two typical open datasets illustrate the superior performance of our approach and the outstanding robustness when appearance changes.
Runwei Ding, Hong Liu Key
ICIP2
2018 Learning Explicit Shape and Motion Evolution Maps for Skeleton-Based Human Action Recognition
abstract
Human action recognition based on skeleton sequences has wide applications in human-computer interaction and intelligent surveillance. Although previous methods have successfully applied Long Short-Term Memory(LSTM) networks to model shape evolution of human actions, it still remains a problem to efficiently recognize actions, especially for similar actions from sequential data due to the lack of the details of motion. To solve this problem, this paper presents an improved LSTM-based network to jointly learn explicit long-term shape evolution maps (SEM) and motion evolution maps (MEM). Firstly, human actions are represented as compact SEM and MEM, which mutually compensate. Secondly, these maps are jointly learned by deep LSTM networks to explore high-level temporal dependencies. Then, a weighted aggregate layer (WAL) is designed to aggregate outputs of L-STM networks cross different temporal stages. Finally, deep features of shape and motion are combined by decision level fusion. Experimental results on the currently largest NTU RGB+D dataset and public SmartHome dataset verify that our method significantly outperforms the state-of-the-arts.
Hong Liu 0008, Juanhui Tu, Runwei Ding
ICASSP4
2018 Audio-Visual Keyword Spotting Based on Multidimensional Convolutional Neural Network
abstract
The fusion of audio and visual information is one of the most promising solutions for reliable keyword spotting (KWS), particularly when audio is corrupted by noise. KWS aims to detect a specific word in an audio stream, which still remains a challenging problem under noisy environments. In this paper, an audio-visual neural network based on multidimensional convolutional neural network (MCNN) is proposed to perform audio-visual KWS. Firstly, the log mel-spectrogram and lip area sequence are extracted, respectively, from the audio and visual streams, and are taken as the input of the audio-visual neural network. Then, an audio-visual neural network based on MCNN consisting of 2D CNN and 3D CNN is used to model the time-frequency feature of the log mel-spectrogram and the spatiotemporal feature of the lip area sequence, respectively. Finally, the outputs of the audio and visual networks are combined for KWS through decision fusion. Experimental results on the PKU-AV database under complex acoustic conditions demonstrate that the proposed method achieves preferable performance compared to other state-of-the-art methods.
Runwei Ding, Hong Liu 0008
ICIP1
2018 Spatial-Temporal Data Augmentation Based on LSTM Autoencoder Network for Skeleton-Based Human Action Recognition
abstract
Data augmentation is known to be of crucial importance for the generalization of RNN-based methods of skeleton-based human action recognition. Traditional data augmentation methods artificially adopt various transformations merely in spatial domain, which lack effective temporal representation. This paper extends traditional Long Short-Term Memory (LSTM) and presents a novel LSTM autoencoder network (LSTM-AE) for spatial-temporal data augmentation. In the LSTM-AE, the LSTM network preserves the temporal information of skeleton sequences, and the autoencoder architecture can automatically eliminate irrelevant and redundant information. Meanwhile, a regularized cross-entropy loss is defined to guide the LSTM-AE to learn more suitable representations of skeleton data. Experimental results on the currently largest NTU RGB+D dataset and public SmartHome dataset verify that the proposed model outperforms the state-of-the-art methods, and can be integrated with most of the RNN-based action recognition models easily.
Juanhui Tu, Hong Liu 0008, Fanyang Meng, Mengyuan Liu 0001, Runwei Ding
ICIP5
2009 Image Restoration of Warped Complex Chinese Documents Based on Text Boundary Lines
abstract
Distortion always appears in document images while scanning thick bound volumes. There are two kinds of distortion for the scanned grayscale images, shadow appears at the volumes' spine area, and warping of the words occurs in the shadow. In this paper, a novel text boundary lines based method for efficient restoration of warped scanning Chinese document images is presented. We first detect on which side of an image the shadow lays by row grayscale analysis method. Then the shadow is removed by a modified Niblack's algorithm. In order to detect the warped feature, a text boundary lines' detection method is proposed. Finally, an adjustment method based on the text boundary lines is carried to restore the warped words. Experiments on 400 various scanning Chinese document images are implemented. The improvement on average character recall is 11.92% to 14.89%. Experiments show that the proposed restoration method is efficient for Chinese documents with both text and non-text regions.
Hong Liu 0008, Runwei Ding
SMC2