Yonghong Hou

dblp:123/2091 · DBLP profile ↗
← Back
58ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0002-1676-5505ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 15 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A multimodal medical segmentation method based on modality complementarity and state space modeling
Yonghong Hou
J. Vis. Commun. Image Represent.2
2026 The triplet loss with angular loss and adaptive margin based metric learning for medical image classification
Yonghong Hou
Mach. Vis. Appl.2
2025 Exemplar-free class incremental action recognition based on self-supervised learning
Chunyu Hou, Yonghong Hou, Jinyin Jiang, Gunel Abdullayeva
Image Vis. Comput.2
2025 EdgeStereoSR: A multi-task network with transformers for stereo image super-resolution considering edge prior
Anqi Liu 0005, Sumei Li, Yongli Chang, Yonghong Hou
Signal Process.4
2025 Visual context learning based on cross-modal knowledge for continuous sign language recognition
Kailin Liu, Yonghong Hou, Zihui Guo
Vis. Comput.2
2025 Semi-supervised medical image segmentation through label-driven space structure augmentation
Zhijun Yan, Yonghong Hou
Vis. Comput.2
2024 Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
abstract
Deep neural networks (DNNs) have been applied in many computer vision tasks and achieved state-of-the-art (SOTA) performance. However, misclassification will occur when DNNs predict adversarial examples which are created by adding human-imperceptible adversarial noise to natural examples. This limits the application of DNN in security-critical fields. In order to enhance the robustness of models, previous research has primarily focused on the unimodal domain, such as image recognition and video understanding. Although multi-modal learning has achieved advanced performance in various tasks, such as action recognition, research on the robustness of RGB-skeleton action recognition models is scarce. In this paper, we systematically investigate how to improve the robustness of RGB-skeleton action recognition models. We initially conducted empirical analysis on the robustness of different modalities and observed that the skeleton modality is more robust than the RGB modality. Motivated by this observation, we propose the Attention-based Modality Reweighter (AMR), which utilizes an attention layer to re-weight the two modalities, enabling the model to learn more robust features. Our AMR is plug-and-play, allowing easy integration with multimodal models. To demonstrate the effectiveness of AMR, we conducted extensive experiments on various datasets. For example, compared to the SOTA methods, AMR exhibits a 43.77% improvement against PGD20 attacks on the NTURGB+D 60 dataset. Furthermore, it effectively balances the differences in robustness between different modalities.
Xin Liu 0012, Zitong Yu, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002
IJCB4
2024 DFN: A deep fusion network for flexible single and multi-modal action recognition
abstract
Multi-modal action recognition methods can be generally classified into two categories: (1) fusing multi-modal features with simple concatenation or fusing the classification scores of individual modalities without considering the interaction among the multi-modalities; (2) using one of the modalities as privileged information in training to boost the recognition on the other modalities in inference. The former approach usually is not able to deal with the cases where one of the modalities is missing. In the latter, the trained classifier does not work on the privileged modality. To address these shortcomings, this paper presents a novel end-to-end trainable deep fusion network (DFN) that is able to improve the performance not only in the cases where all modalities are available and also in the cases where there is a missing modality. The DFN is simple yet effective with the capability of retrieving an estimation of one modality by using another modality through a Multilayer Perceptron (MLP). In order to better preserve structure information, the DFN first maps the individual modality features to a high dimensional Kronecker-product space and subsequently learns a low-dimensional discriminative space for classification. The effectiveness of the proposed DFN has been verified on three benchmark datasets: the large NTU RGB+D, UTD-MHAD, and SYSU-3D datasets and it has achieved state-of-the-art results.
Chuankun Li, Yonghong Hou, Wanqing Li 0001, Zewei Ding, Pichao Wang
Expert Syst. Appl.2
2024 Dual states based reinforcement learning for fast MR scan and image reconstruction
Yanwei Pang, Xuebin Sun, Yonghong Hou, Zhenghan Yang, Zhenchang Wang
Neurocomputing4
2024 FTAN: Frame-to-frame temporal alignment network with contrastive learning for few-shot action recognition
Yonghong Hou, Zihui Guo, Zhiyi Gao
Image Vis. Comput.2
2024 GNSS-R Ocean Wind Speed Retrieval Algorithm Based on Fusing Frequency-Domain Information
abstract
Ocean surface wind speed is important for numerical prediction of the marine environment, marine disaster monitoring, sea–steam interaction, meteorological prediction, climate research, and so on. At present, ocean surface wind speed retrieval models extract DDM features from the delay-Doppler domain, the angle of which is single, where certain detail information cannot be effectively extracted from the image. To improve the accuracy of wind speed retrieval, starting from the feature extraction, this letter proposes a frequency-informed neural network (FINN)-based wind speed retrieval. First, the wind speed retrieval model based on the delay-Doppler domain and frequency domain is constructed, based on the extraction of features from the DDM delay-Doppler domain, the features from the DDM frequency domain are also extracted, so the wind speed retrieval is performed separately by extracting features from different angles. Then, a multimodel fusion scheme based on dynamic weights is designed, which can dynamically weight submodels according to different input samples, and realize the dynamic fusion of delay-Doppler-domain retrieval results and frequency-domain retrieval results. Finally, the improved gradient loss (IG Loss) function is proposed. Contrast experiments and ablation studies prove that the present algorithm has excellent retrieval performance.
Hongchen Liu, Yonghong Hou, Shuang Jiang, Meiyan Huang, Hongbo Qu
IEEE Geosci. Remote. Sens. Lett.2
2024 Global and Local Contrastive Learning for Self-Supervised Skeleton-Based Action Recognition
abstract
Contrastive learning for self-supervised skeleton-based action recognition has recently received attention. It has been observed that local crops, containing partial action sequences, can predict action patterns, which is advantageous for skeleton-based action recognition. This paper proposes a Global and Local Contrastive Learning framework (skeleton-logoCLR) with two contrastive learning routes, Global-to-Global and Global-to-Local, which utilize the similarity between global and local crops of the same skeleton sequence. Specifically, in the Global-to-Global route, we design Temporal Attention Crop-Resize (TACR) to learn global semantic information by maximizing the retention of action region in the temporal dimension. In the Global-to-Local route, the proposed Skeleton-logo Augmentation is deviced to concatenate two local crops from different sequences for local semantic learning. Moreover, instead of fusing directly, the losses of two routes are combined in a cascaded manner through the Self-Adaptive Training Strategy (SATS) to achieve stronger generalization performance. Extensive experiments are conducted on the NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets. The results demonstrate that the proposed method achieves remarkable performance.
Jinhua Hu, Yonghong Hou, Zihui Guo, Jiajun Gao
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multi-Scale Visual Perception Based Progressive Feature Interaction Network for Stereo Image Super-Resolution
abstract
In recent years, stereo image super-resolution based on convolutional neural network has been extensively researched and achieved impressive performance by introducing complementary information from another view. However, most existing methods still cannot fully capture both intra- and cross-view information due to the neglect of multi-scale information perception, multi-scale binocular alignment and the excitation of large scale to small scale in human vision system. And they generated blurry results due to the consideration of irrelevant information in search for cross-view information. To address these issues, we propose a multi-scale visual perception based progressive feature interaction network (MS-PFINet) for stereo image super-resolution. Specifically, to exploit comprehensive intra- and cross-view information for image reconstruction, we design a two-stream network with multi-branch structure to extract multi-scale features and progressively use cross-view interaction at larger scales to guide that at smaller scales. Moreover, to explore more proper and accurate cross-view information, we propose a feature transformer module (FTM) to search and transfer the most relevant features from another view by hard attention maps and soft attention maps, which are calculated by patch-wise similarity rather than pixel-wise. In addition, in order to encourage a more effective way to transfer texture features for the target view, we propose a perceptual texture matching loss to supervise the accuracy of feature transformer modules. Experimental results show that our proposed method is superior to the state-of-the-art methods in most cases.
Anqi Liu 0005, Sumei Li, Yongli Chang, Yonghong Hou
IEEE Trans. Circuits Syst. Video Technol.4
2024 Spatial-Temporal Enhanced Network for Continuous Sign Language Recognition
abstract
Continuous Sign language Recognition (CSLR) aims to generate gloss sequences based on untrimmed sign videos. Since discriminative visual features are essential for CSLR, current efforts mainly focus on strengthening the feature extractor. The feature extractor can be disassembled into a spatial representation module and a short-term temporal module for spatial and visual features modeling. However, existing methods always regard it as a monoblock and rarely implement specific refinements for such two distinct modules, which is difficult to achieve effective modeling of spatial appearance information and temporal motion information. To address the above issues, we proposed a spatial temporal enhanced network which contains a spatial-visual alignment (SVA) module and a temporal feature difference (TFD) module. Specifically, the SVA module conducts an auxiliary task between the spatial features and target gloss sequences to enhance the extraction of hand and facial expressions. Meanwhile, the TFD module is constructed to exploit the underlying dynamic between consecutive frames and inject the aggregated motion information into spatial features to assist short-term temporal modeling. Extensive experimental results demonstrate the effectiveness of the proposed modules and our network achieves state-of-the-art or competitive performance on four public CSLR datasets.
Yonghong Hou, Zihui Guo, Kailin Liu
IEEE Trans. Circuits Syst. Video Technol.2
2024 SWGNet: Step-Wise Reference Frame Generation Network for Multiview Video Coding
abstract
In multiview video coding, the coding performance highly depends on the quality of the reference frames. In view of this, a step-wise reference frame generation network (SWGNet) is designed to improve the quality of the reference frame for efficient multiview video coding. In particular, a frame-level to block-level learning paradigm is proposed to step-wisely generate a high-quality reference frame. In the frame-level stage, by exploiting parallax correlations between temporal and inter-view references on the basis of image alignment, a parallax-guided frame-level synthesis module is proposed to generate an elementary reference frame. Then, in the block-level stage, a transformer-based block-level aggregation module is designed to further refine the texture details of the reference frame by modeling long-range dependencies among pixels. The proposed SWGNet is integrated into 3D-HEVC, and extensive experiments demonstrate that the proposed method achieves significant bitrate saving compared with 3D-HEVC.
Jing Zhang 0017, Yonghong Hou, Zhaoqing Pan, Bo Peng 0007, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 WSPTGAN for Global Ocean Surface Wind Speed Generation With High Temporal Resolution and Spatial Coverage
abstract
Obtaining global ocean surface wind speed data with high temporal resolution and spatial coverage is a challenging task. Due to the lack of widely applicable direct measurement methods and algorithms, current research and data products can only achieve good performance in a small spatial range or at low temporal resolution. In this article, a generative adversarial network (GAN) with a transformer structure called Wind Speed Prediction transformer-GAN (WSPTGAN) is proposed to generate wind speed data with good spatial coverage and high temporal resolution for areas. The WSPTGAN is trained with the proposed image-like wind speed data combined partial missing dataset (CPMD), which is combined with the fifth generation of the European Center for Medium-Range Weather Forecast (ECMWF) reanalysis data and Advanced Scatterometer (ASCAT) data from Meteorological Operational satellites. Thanks to the defective data learning mechanism (DDLM), sequential-wise multihead self-attention mechanism (SMSM), and sequence feature adaptive verification mechanism (SFAVM) in the proposed algorithm, the obtained model has good wind speed prediction accuracy with root mean square error (RMSE) of 0.8984 m/s and can achieve multistep 10-min wind speed data generation within the global ocean. After comparison with five state-of-the-art prediction models, it is confirmed that the algorithm in this article is able to make better use of the defective data for learning and prediction of wind field trends in global ocean regions.
Yonghong Hou, Xiaowei Song 0001, Chunping Hou, Zixiang Xiong, Dan Ma 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 Self-Attention-Guided Multiindicator Retrieval for Ocean Surface Wind Field With Multimodal Data Augmentation and Fusion
abstract
The deployment of global navigation satellite system reflectometry (GNSS-R) emerges as a compelling approach for the extraction of ocean surface wind field, primarily due to its exceptional cost-effectiveness, all-weather robustness, and excellent spatiotemporal coverage. Despite these advantages, the insufficient use of various data and the lack of ability to perform multiindicator retrieval limit the performance of existing methods in practical ocean wind field retrieval. To overcome these limitations, this article introduces a novel self-attention-guided ocean surface wind field multiindicator retrieval algorithm based on multimodal observation data augmentation and fusion. Initially, data generation modules are employed to complement high-quality observation data that are not fully provided by GNSS-R system. Subsequently, the multiscale data fusion encoder (MDFE) is implemented to extract and fuse multiple data features of different scales to enhance the utilization ability of data and improve the accuracy of wind field retrieval. Finally, the self-attention multiindicator predictor (SAMIP) is put to use for optimizing the feature attention strategy, achieving accurate retrieval of ocean surface wind speed and direction simultaneously. The proposed method provides a novel solution for the comprehensive utilization of various data products in the GNSS-R system, simultaneously achieves synchronous retrieval of multiple wind field indicators, which are wind speed and direction. In the context of recent advancements in algorithmic development, the proposed algorithm exhibits a notable enhancement in the precision of wind speed and direction retrieval. Benchmarking against ERA5 wind field data, the proposed algorithm achieves a root-mean-square error (RMSE) of 1.23 m/s in wind speed retrieval, demonstrating at least a 9% accuracy improvement compared to five state-of-the-art algorithms from recent years. Furthermore, the RMSE in wind direction retrieval stands at 20.7°, surpassing comparison algorithms by achieving a reduction of over 8% in error. Collectively, these metrics robustly validate the efficacy of the proposed algorithm.
Yonghong Hou, Xiaowei Song 0001, Chunping Hou, Zixiang Xiong, Dan Ma 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 Image Reconstruction for Accelerated MR Scan With Faster Fourier Convolutional Neural Networks
abstract
High quality image reconstruction from undersampled k -space data is key to accelerating MR scanning. Current deep learning methods are limited by the small receptive fields in reconstruction networks, which restrict the exploitation of long-range information, and impede the mitigation of full-image artifacts, particularly in 3D reconstruction tasks. Additionally, the substantial computational demands of 3D reconstruction considerably hinder advancements in related fields. To tackle these challenges, we propose the following: 1) A novel convolution operator named Faster Fourier Convolution (FasterFC), aims at providing an adaptable broad receptive field for spatial domain reconstruction networks with fast computational speed. 2) A split-slice strategy that substantially reduces the computational load of 3D reconstruction, enabling high-resolution, multi-coil, 3D MR image reconstruction while fully utilizing inter-layer and intra-layer information. 3) A single-to-group algorithm that efficiently utilizes scan-specific and data-driven priors to enhance k -space interpolation effects. 4) A multi-stage, multi-coil, 3D fast MRI method, called the faster Fourier convolution based single-to-group network (FAS-Net), comprising a single-to-group k -space interpolation algorithm and a FasterFC-based image domain reconstruction module, significantly minimizes the computational demands of 3D reconstruction through split-slice strategy. Experimental evaluations conducted on the NYU fastMRI and Stanford MRI Data datasets reveal that the FasterFC significantly enhances the quality of both 2D and 3D reconstruction results. Moreover, FAS-Net, characterized as a method that can achieve high-resolution (320, 320, 256), multi-coil, (8 coils), 3D fast MRI, exhibits superior reconstruction performance compared to other state-of-the-art 2D and 3D methods.
Yanwei Pang, Xuebin Sun, Yonghong Hou, Zhenchang Wang, Xuelong Li 0001
IEEE Trans. Image Process.5
2024 Coarse-to-Fine Cross-View Interaction Based Accurate Stereo Image Super-Resolution Network
abstract
Recently, parallax attention based stereo image super-resolution (SR) methods, which can better explore cross-view information, have been widely studied. Despite the impressive performance of these methods, almost all of them calculate parallax attention maps at a single low resolution, which will lead to ambiguous stereo correspondence. Besides, the widely used parallax attention module (PAM) cannot handle the illuminance variations in stereo image pairs, and cannot distinguish the contribution of the captured cross-view features to the reconstruction of the target view. To this end, in this paper, we propose a coarse-to-fine cross-view interaction based network (C2FNet) to achieve more accurate cross-view information capturing. Firstly, in C2FNet, a coarse-to-fine cascaded parallax attention structure (C2F-CPAS), which conforms with the human visual mechanism, is constructed to gradually perform parallax attention from the low-resolution to high-resolution level. Thus, richer textures can be used to learn more reliable stereo correspondence. Meanwhile, a multi-level attention transfer loss is designed to further calibrate the accuracy of stereo correspondence at each level. Secondly, we propose a modified PAM (MPAM) to alleviate the limitations of common PAM so that illuminance-robust stereo correspondence can be learned and more important cross-view information can be selected. Extensive experimental results show that our proposed C2FNet outperforms the state-of-the-art methods on various datasets.
Anqi Liu 0005, Sumei Li, Yongli Chang, Yonghong Hou
IEEE Trans. Multim.5
2023 Discovery the inverse variational problems from noisy data by physics-constrained machine learning
Hongbo Qu, Hongchen Liu, Shuang Jiang, Yonghong Hou
Appl. Intell.5
2023 EVOLVE: Learning volume-adaptive phases for fast 3D magnetic resonance scan and image reconstruction
Yanwei Pang, Xuebin Sun, Yonghong Hou
Neurocomputing4
2023 Motion saliency based hierarchical attention network for action recognition
Zihui Guo, Yonghong Hou, Renyi Xiao, Chuankun Li, Wanqing Li 0001
Multim. Tools Appl.2
2023 Sign language recognition via dimensional global-local shift and cross-scale aggregation
Zihui Guo, Yonghong Hou, Wanqing Li 0001
Neural Comput. Appl.2
2023 FT-HID: a large-scale RGB-D dataset for first- and third-person human interaction analysis
Zihui Guo, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001
Neural Comput. Appl.2
2023 Locality-Aware Transformer for Video-Based Sign Language Translation
abstract
Recently, the application of transformer makes significant progress in sign language translation. However, several characteristics of sign videos are neglected in existing transformer-based methods that hinder translation performance. Firstly, in sign videos, multiple consecutive frames represent a single sign gloss thus the local temporal relations are crucial. Secondly, the inconsistency between video and text demands the non-local and global context modeling ability of the model. To address these issues, a locality-aware transformer is proposed for sign language translation. Concretely, the multi-stride position encoding scheme assigns the same position index to adjacent frames with various strides to enhance the local dependency. Afterward, the adaptive temporal interaction module is utilized to capture non-local and flexible local frame correlation simultaneously. Moreover, a gloss counting task is designed to facilitate the holistic understanding of sign videos. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed framework.
Zihui Guo, Yonghong Hou, Chunping Hou
IEEE Signal Process. Lett.2
2023 Learning Spatio-Temporal Semantics and Cluster Relation for Zero-Shot Action Recognition
abstract
Zero-shot Action Recognition (ZSAR) aims at bridging the video$\rightarrow $class relation with only labeled training data of seen classes while generalizing the model to alleviate the heterogeneity of unseen actions. Most existing methods have comprehensively represented videos and action classes, however, the semantic gap and the hubness problem between them remain crucial challenges that are under-explored. In this paper, we propose an effective method to tackle the above issues. Specifically, to narrow the semantic gap, we end-to-end generate a spatio-temporal semantics for each video, which provides essential textual information to refine the video representation. Furthermore, we propose a compactness-separability loss that optimizes the intra- and inter-class relations in a unified formula and quantitatively constrains cluster distribution, thus effectively diminishing the impact of the hubness problem. Extensive experiments on UCF101, HMDB51, and Olympic Sports datasets prove the effectiveness of the proposed approach and demonstrate our approach outperforms the state-of-the-art methods.
Jiajun Gao, Yonghong Hou, Zihui Guo, Haochun Zheng
IEEE Trans. Circuits Syst. Video Technol.2
2023 S2G2HAD: A Graph-Guided Siamese Reconstruction Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to identify anomalous pixels in the image with significant spectral differences from their surrounding background pixels, and has important military and civilian applications. However, challenges such as high spectral similarity between pixels, data redundancy, and lack of prior information pose significant difficulties in HAD. To address these issues, we propose the Selective Siamese Graph Guided Hyperspectral Anomaly Detection method. In the initial phase of this study, we present a spatial-spectral feature dynamic composition module. Guided by the principles of three-way clustering theories, this module excels in achieving heightened precision in feature embedding and graph construction. Subsequently, we introduce the Characteristic Expression Differentiation mechanism, designed to enhance the separation of hidden layer features for heterogeneous data in high-dimensional space by incorporating Siamese network-derived similarity discrimination principles. Lastly, we develop a graph-guided Siamese selective reconstruction module that places a strong emphasis on the differentiation between background and anomaly features. It utilizes background graph data to construct a pure background and concurrently establishes connections across spectral bands. This approach significantly enhances data processing efficiency while reducing computational resource consumption through the elimination of redundancy. Extensive experiments on seven public datasets demonstrate the effectiveness of the method.
Dan Ma 0003, Yonghong Hou, Beichen Li 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Learning Using Privileged Information for Zero-Shot Action Recognition
Zhiyi Gao, Yonghong Hou, Wanqing Li 0001, Zihui Guo
ACCV (4)2
2022 Contrastive Positive Mining for Unsupervised 3D Action Representation Learning
Yonghong Hou, Wanqing Li 0001
ECCV (4)2
2022 Skeletal Twins: Unsupervised Skeleton-Based Action Representation Learning
abstract
In this paper, we investigate unsupervised representation learning for skeleton action recognition, and develop a simple yet effective framework: SKeletal Twins (SKT), which is capable of learning representations from unlabeled skeleton data. To be specific, we choose skeleton-specific spatial and temporal augmentations for spatio-temporal dynamics learning, then the augmented skeleton sequence is represented as a graph with both spatial and temporal edges so that the GCN-based twin encoders are able to encode human pose and joint's temporal motion. Barlow Twins' objective function is used to minimize the redundancy and keep similarity of different skeleton augmentations. However it ignores the instance-level consistency of the skeleton instance from different augmentations, thus an instance-level consistency-enhanced objective function is designed and jointly optimized, which boosts the representation learning. Extensive experiments verify that the proposed framework obtains the state-of-the-art results on the challenging NTU-60 and NTU-120 datasets.
Yonghong Hou
ICME2
2022 Corrigendum: User-Experience-Oriented Fuzzy Logic Controller for Adaptive Streaming
abstract
doi:10.1093/comjnl/bxy010 Comp J 2018;61(7): 1064–1074 The affiliations of the following authors have been corrected since the original publication of this article: Yonghong Hou Lin Xue Shuo Li Jiaming Xing
Yonghong Hou, Shuo Li 0003, Jiaming Xing
Comput. J.1
2022 Deep region segmentation-based intra prediction for depth video coding
Jing Zhang 0017, Yonghong Hou, Zhe Zhang 0041, Dengchao Jin, Peihan Zhang, Ge Li 0002
Multim. Tools Appl.2
2022 Unsupervised skeleton-based action representation learning via relation consistency pursuit
Yonghong Hou
Neural Comput. Appl.2
2022 A Central Difference Graph Convolutional Operator for Skeleton-Based Action Recognition
abstract
This paper proposes a new graph convolutional operator called central difference graph convolution (CDGC) for skeleton based action recognition. It is not only able to aggregate node information like a vanilla graph convolutional operation but also gradient information. Without introducing any additional parameters, CDGC can replace vanilla graph convolution in any existing Graph Convolutional Networks (GCNs). In addition, an accelerated version of the CDGC is developed which greatly improves the speed of training. Experiments on two popular large-scale datasets NTU RGB+D 60 & 120 have demonstrated the efficacy of the proposed CDGC. Code is available athttps://github.com/iesymiao/CD-GCN.
Shuangyan Miao, Yonghong Hou, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Transformer guided geometry model for flow-based unsupervised visual odometry
Xiangyu Li 0009, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001
Neural Comput. Appl.2
2021 Deep edge map guided depth super resolution
Zhongyu Jiang, Huanjing Yue, Yukun Lai, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou
Signal Process. Image Commun.5
2021 Deep noise estimation and removal for real-world noisy images
Huanjing Yue, Zhongyu Jiang, Shengdi Zhou, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou
Signal Process. Image Commun.5
2020 SAR-NAS: Skeleton-based action recognition via neural architecture searching
Yonghong Hou, Pichao Wang, Zihui Guo, Wanqing Li 0001
J. Vis. Commun. Image Represent.2
2020 ConvNets-based action recognition from skeleton motion maps
Yanfang Chen, Chuankun Li, Yonghong Hou, Wanqing Li 0001
Multim. Tools Appl.4
2019 Single Image De-Raining via Generative Adversarial Nets
abstract
In this paper, we propose a Generative Adversarial Network for Single Image De-raining(GAN-SID). We observe that batch normalization has side effects in the de-raining task. Therefore, we introduce instance normalization to replace the traditional batch normalization layers in both generator and discriminator. Motivated by the Squeeze-and-Excitation (SE) network that can learn the importance of channels, we introduce SE module in the generator to give different weights to the learned features. To preserve image details while removing rain streaks, we propose to utilize pixel-wise loss, perceptual loss, and adversarial loss to train the proposed network. Experiments on two synthetic datasets and real world images demonstrate that the proposed method outperforms state-of-the-art de-raining works in both objective and subjective measurements.
Shichao Li 0006, Yonghong Hou, Huanjing Yue, Zihui Guo
ICME2
2019 Self-Attention Guided Deep Features for Action Recognition
abstract
Skeleton based human action recognition is an important task in computer vision. However, it is very challenging due to the complex spatio-temporal variations of skeleton joints. In this work, we propose an end-to-end trainable network consisting of a Deep Convolutional Model (DCM) and a Self-Attention Model (SAM) for human action recognition from skeleton data. Specifically, skeleton sequences are encoded into color images and fed into DCM to extract deep features. In the SAM, handcrafted features representing the motion degree of joints are extracted and the attention weights are learned by a simple yet effective linear mapping. The effectiveness of proposed method has been verified on NTU RGB+D, SYSU-3D and UTD-MHAD datasets and achieved state-of-the-art results.
Renyi Xiao, Yonghong Hou, Zihui Guo, Chuankun Li, Pichao Wang, Wanqing Li 0001
ICME2
2019 Light Weight Stereo Matching via Deep Extraction and Integration of Low and High Level Information
abstract
Deep convolutional neural networks (CNN) have demonstrated remarkable progress in stereo matching recently. However, disparity estimation in the ill-posed regions is still difficult. In addition, CNN based stereo matching methods often have impractical computational complexity and memory consumption. To address these problems we propose an end-to-end light weight CNN architecture to effectively learn and integrate low and high level information. To achieve this, a novel enhancement block built upon group convolution and dilated-convolution is proposed. Compared with state-of-the-art methods, the proposed method achieved competitive performance with the least number of network parameters on the Flyingthings3d and KITTI datasets.
Yonghong Hou, Pichao Wang, Zhongyu Jiang, Wanqing Li 0001
ICME2
2019 DVONet: Unsupervised Monocular Depth Estimation and Visual Odometry
abstract
This paper proposes an unsupervised learning framework for monocular depth estimation and visual odometry (VO), referred to as DVONet. The framework is trained using stereo image sequences and is able to estimate absolute-scale scene depth and camera poses from monocular images. To mitigate the effect of stereo occlusions in training and improve the depth estimation, left-right occlusion mask is introduced. In addition, a novel VO network is proposed where the feature extraction network is shared between pose estimation and optical flow estimation. The proposed DVONet achieves state-of-the-art results for both depth estimation and VO tasks on the KITTI driving dataset, outperforming the existing unsupervised methods and being comparable to the traditional ones.
Xiangyu Li 0009, Yonghong Hou, Pichao Wang, Wanqing Li 0001
VCIP2
2019 Learning attentive dynamic maps (ADMs) for Understanding Human Actions
Chuankun Li, Yonghong Hou, Wanqing Li 0001, Pichao Wang
J. Vis. Commun. Image Represent.2
2019 Multiview-Based 3-D Action Recognition Using Deep Networks
abstract
In multiview learning, views may be obtained from multiple sources or extracted from a single source as different features. In this paper, effective multiple views from skeleton sequences are proposed to learn the discriminative features using multiple networks for three-dimensional human action recognition. Specifically, three views are constructed in the spatial domain and fed to a stack of long short-term memory networks to exploit temporal information and three views are constructed using the improved joint trajectory maps and fed to three convolutional neural networks to exploit spatial information. Multiply fusion is used to combine the recognition scores of all views. The proposed method has been verified and achieved the state-of-the-art results on the widely used UTD-MHAD, MSRC-12 Kinect Gesture, and NTU red, green, blue (RGB)+D datasets.
Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Trans. Hum. Mach. Syst.2
2018 User-Experience-Oriented Fuzzy Logic Controller for Adaptive Streaming
abstract
HyperText Transfer Protocol (HTTP) streaming has been widely used for multimedia delivery nowadays. To adapt to the heterogeneous networks and different terminals, the rate adaptation controller is developed as the core part of the video transmission system. In this paper, a three-input fuzzy controller is designed to enhance the users’ quality-of-experience (QoE) on watching online video. The introduction of fuzzy logic is able to solve the problem that there is no rule to set the reasonable buffer thresholds in the buffer-based adaptation controller. The normalized throughput, the buffer level and the buffer variation are used as the inputs of the controller to minimize the inevitable mismatch caused by limited bitrate levels and to help the system converge to a stable state. The proposed controller is compared with three other controllers under three simulation network conditions and an actual vehicle condition. A QoE model is simultaneously used to get intuitive and comprehensive results. The results show that among all the four conditions, the proposed controller can provide better QoE than other algorithms.
Yonghong Hou, Shuo Li 0003, Jiaming Xing
Comput. J.1
2018 Action recognition based on joint trajectory maps with convolutional neural networks
Pichao Wang, Wanqing Li 0001, Chuankun Li, Yonghong Hou
Knowl. Based Syst.4
2018 Combining ConvNets with hand-crafted features for action recognition based on an HMM-SVM classifier
Shuang Wang 0001, Yonghong Hou, Jiarong Dong, Chang Tang
Multim. Tools Appl.2
2018 Skeleton Optical Spectra-Based Action Recognition Using Convolutional Neural Networks
abstract
This letter presents an effective method to encode the spatiotemporal information of a skeleton sequence into color texture images, referred to as skeleton optical spectra, and employs convolutional neural networks (ConvNets) to learn the discriminative features for action recognition. Such spectrum representation makes it possible to use a standard ConvNet architecture to learn suitable “dynamic” features from skeleton sequences without training millions of parameters afresh and it is especially valuable when there is insufficient annotated training video data. Specifically, the encoding consists of four steps: mapping of joint distribution, spectrum coding of joint trajectories, spectrum coding of body parts, and joint velocity weighted saturation and brightness. Experimental results on three widely used datasets have demonstrated the efficacy of the proposed method.
Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Depth Super-Resolution From RGB-D Pairs With Transform and Spatial Domain Regularization
abstract
This paper proposes a depth super-resolution method with both transform and spatial domain regularization. In the transform domain regularization, nonlocal correlations are exploited via an auto-regressive model, where each patch is further sparsified with a locally-trained transform to consider intra-patch correlations. In the spatial domain regularization, we propose a multi-directional total variation (MTV) prior to characterize the geometrical structures spatially orientated at arbitrary directions in depth maps. To achieve adaptive regularization, the MTV is weighted for each directional finite difference considering local characteristics of RGB-D data. We develop an accelerated proximal gradient algorithm to solve the proposed model. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior depth super-resolution performance for various configurations of magnification factors and datasets.
Zhongyu Jiang, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002, Chunping Hou
IEEE Trans. Image Process.2
2017 Sketch based image retrieval via image-aided cross domain learning
abstract
Existing methods on sketch based image retrieval (SBIR) are usually based on the hand-crafted features whose ability of representation is limited. In this paper, we propose a sketch based image retrieval method via image-aided cross domain learning. First, the deep learning model is introduced to learn the discriminative features. However, it needs a large number of images to train the deep model, which is not suitable for the sketch images. Thus, we propose to extend the sketch training images via introducing the real images. Specifically, we initialize the deep models with extra image data, and then extract the generalized boundary from real images as the sketch approximation. The using of generalized boundary is under the assumption that their domain is similar with sketch domain. Finally, the neural network is fine-tuned with the sketch approximation data. Experimental results on Flicker15 show that the proposed method has a strong ability to link the associated image-sketch pairs and the results outperform state-of-the-arts methods.
Jianjun Lei 0001, Kaifu Zheng, Hua Zhang 0008, Xiaochun Cao, Nam Ling, Yonghong Hou
ICIP6
2017 Rate control for HEVC based on spatio-temporal context and motion complexity
Yonghong Hou, Jianjun Lei 0001, Wei Xiang 0001, Yao Guo 0006
Multim. Tools Appl.1
2017 Joint Distance Maps Based Action Recognition With Convolutional Neural Networks
abstract
Motivated by the promising performance achieved by deep learning, an effective yet simple method is proposed to encode the spatio-temporal information of skeleton sequences into color texture images, referred to as joint distance maps (JDMs), and convolutional neural networks are employed to exploit the discriminative features from the JDMs for human action and interaction recognition. The pair-wise distances between joints over a sequence of single or multiple person skeletons are encoded into color variations to capture temporal information. The efficacy of the proposed method has been verified by the state-of-the-art results on the large RGB+D Dataset and small UTD-MHAD Dataset in both single-view and cross-view settings.
Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Signal Process. Lett.2
2017 Near-Optimal Cross-Layer Forward Error Correction Using Raptor and RCPC Codes for Prioritized Video Transmission Over Wireless Channels
abstract
Cross-layer forward error correction (FEC) aims at utilizing available bandwidth more efficiently, which has been applied to error-prone video transmission over imperfect wireless channels. In this paper we propose a new near-optimal cross-layer FEC scheme in which systematic Raptor codes are used at the application layer and rate compatible punctured convolutional (RCPC) codes are used at the physical layer for H.264/AVC encoded video streaming with channel bandwidth constraints. In the proposed scheme, in order to fully exploit the unequal importance of compressed video data, we assign each source packet a different priority according to its contribution to the reconstructed video quality. We first obtain the transmission parameters, which satisfies the conditions for optimal video transmission, for the optimal cross-layer Raptor-RCPC FEC in the ideal situation through a theoretical analysis, and then we propose a heuristic algorithm searching from the optimal solution point to obtain the transmission parameters, which are near optimal in the practical situation. Computer simulation results show that the proposed scheme can achieve significant performance improvements in both the additive white Gaussian noise and Rayleigh channels compared with the previous work.
Yonghong Hou, Wei Xiang 0001, Maode Ma, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.1
2016 Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural Networks
abstract
Recently, Convolutional Neural Networks (ConvNets) have shown promising performances in many computer vision tasks, especially image-based recognition. How to effectively use ConvNets for video-based recognition is still an open problem. In this paper, we propose a compact, effective yet simple method to encode spatio-temporal information carried in 3D skeleton sequences into multiple 2D images, referred to as Joint Trajectory Maps (JTM), and ConvNets are adopted to exploit the discriminative features for real-time human action recognition. The proposed method has been evaluated on three public benchmarks, i.e., MSRC-12 Kinect gesture dataset (MSRC-12), G3D dataset and UTD multimodal human action dataset (UTD-MHAD) and achieved the state-of-the-art results.
Pichao Wang, Yonghong Hou, Wanqing Li 0001
ACM Multimedia3
2016 A Spectral and Spatial Approach of Coarse-to-Fine Blurred Image Region Detection
abstract
Blur exists in many digital images, it can be mainly categorized into two classes: defocus blur which is caused by optical imaging systems and motion blur which is caused by the relative motion between camera and scene objects. In this letter, we propose a simple yet effective automatic blurred image region detection method. Based on the observation that blur attenuates high-frequency components of an image, we present a blur metric based on the log averaged spectrum residual to get a coarse blur map. Then, a novel iterative updating mechanism is proposed to refine the blur map from coarse to fine by exploiting the intrinsic relevance of similar neighbor image regions. The proposed iterative updating mechanism can partially resolve the problem of differentiating an in-focus smooth region and a blurred smooth region. In addition, our iterative updating mechanism can be integrated into other image blurred region detection algorithms to refine the final results. Both quantitative and qualitative experimental results demonstrate that our proposed method is more reliable and efficient compared to various state-of-the-art methods.
Chang Tang, Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Signal Process. Lett.3
2015 A depth estimating method from a single image using FoE CRF
Chunping Hou, Liangzhou Pu, Yonghong Hou
Multim. Tools Appl.4
2013 Performance Optimization of Digital Spectrum Analyzer With Gaussian Input Signal
abstract
Analog to digital converters (ADC) and cascade integrator-comb (CIC) filters are the basic modules in a digital intermediate frequency (IF) spectrum analyzer. The optimal output signal-to-noise ratio (SNR) of the digital IF spectrum analyzer with the Gaussian input signal is considered in this letter. The idea is to strike a trade-off between the saturation error and granular error when quantizing the Gaussian input signal. This letter firstly derives a relationship among the maximum allowed input signal amplitude, input signal power, ADC quantization bits and optimal quantization SNR. Besides, an optimal clipping strategy for the CIC decimation filter with variable decimation rates is proposed. Both numerical and simulation results are presented to demonstrate that the proposed clipping method is able to achieve significant SNR gain compared with the traditional rounding or truncation method.
Yonghong Hou, Guihua Liu, Qing Wang 0015, Wei Xiang 0001
IEEE Signal Process. Lett.1