Yong Zhao 0010

dblp:14/1356-10 · DBLP profile ↗
← Back
53ranked-venue papers
0as first author
22since 2021 · last 2025
0000-0002-7999-1083ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 16 since 2021Artificial intelligence and machine learning · 19 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Rethinking Generalized Long-Tailed Classification: A Constrained Optimization Formulation and Exact Penalty Method
abstract
Long-tail distribution includes two kinds of imbalance: the class-wise imbalance and the attribute-wise one. Sufficient studies have been conducted on the class-wise imbalance. However, the distribution with both the class-wise and the attribute-wise imbalance is more difficult since the attribute is unknown. To solve this problem, we introduce a novel Generalized Long-tail Exact Penalty Method (GL-EPM). GLEPM redefines the long-tail classification as a constrained optimization problem where the model’s prediction on all samples, both the head samples and the tail samples, are constrained to ally with their labels. We then tackles the constrained problem using the exact penalty method, transforming the problem into a series of unconstrained penalty problems. Each penalty problem in the series is generated by the solution of the previous one and the solution series can be theoretically proved to converge to the solution of the constrained optimization problem. We further solve the penalty problems with a rapid-update solution, where the model is trained with dynamically updated sampling rates and loss weights. To validate the efficacy of our proposed method, we conduct a fair comparison with existing baselines across all generalized long-tail benchmarks, demonstrating state-of-the-art performance.
Canmiao Fu, Yong Zhao 0010, Xinan Wang
IJCNN3
2025 A frequency-domain dynamic amplitude filtering method for single-image dehazing with harmony enhancement
Yabo Wu, Yongjun Zhang 0007, Ziyang Chen 0002, Yong Zhao 0010
Expert Syst. Appl.4
2024 Efficient Fusion of Depth Information for Defocus Deblurring
abstract
Defocus deblurring is a classic problem in image restoration tasks. The formation of its defocus blur is related to depth. Recently, the use of dual-pixel sensor designed according to depth-disparity characteristics has brought great improvements to the defocus deblurring task. However, the difficulty of real-time acquisition of dual-pixel images brings difficulties to algorithm deployment. This inspires us to remove defocus blur by single image with depth information. We propose a single-image depth-enhanced defocus deblurring network, which uses a depth map estimated by the monocular depth estimation network to guide the network defocus deblurring. We design a deep information fusion unit, which greatly improves the effect of deblurring. Experiments show that on the single image defocus deblurring task, the experimental results demonstrate the superiority of our method.
Jucai Zhai, Pengcheng Zeng, Chihao Ma, Xinan Wang, Yong Zhao 0010
ICASSP6
2023 Learnable Blur Kernel for Single-Image Defocus Deblurring in the Wild
abstract
Recent research showed that the dual-pixel sensor has made great progress in defocus map estimation and image defocus deblurring. However, extracting real-time dual-pixel views is troublesome and complex in algorithm deployment. Moreover, the deblurred image generated by the defocus deblurring network lacks high-frequency details, which is unsatisfactory in human perception. To overcome this issue, we propose a novel defocus deblurring method that uses the guidance of the defocus map to implement image deblurring. The proposed method consists of a learnable blur kernel to estimate the defocus map, which is an unsupervised method, and a single-image defocus deblurring generative adversarial network (DefocusGAN) for the first time. The proposed network can learn the deblurring of different regions and recover realistic details. We propose a defocus adversarial loss to guide this training process. Competitive experimental results confirm that with a learnable blur kernel, the generated defocus map can achieve results comparable to supervised methods. In the single-image defocus deblurring task, the proposed method achieves state-of-the-art results, especially significant improvements in perceptual quality, where PSNR reaches 25.56 dB and LPIPS reaches 0.111.
Jucai Zhai, Pengcheng Zeng, Chihao Ma, Yong Zhao 0010
AAAI5
2023 CLIP4Stereo: Revisiting Domain Generalized Stereo Matching via Clip
abstract
Despite supervised deep stereo matching networks have achieved impressive performance given sufficient training data, the poor generalization ability caused by the domain shifts prevents them from being applied to unseen domains. Recent progress has shown that CLIP could be a promising alternative for zero-shot visual representation task under the natural language supervision. In this paper, we present a new framework for domain generalized stereo matching by leveraging the contrastive language-image pre-training (CLIP), which distills text-guided discriminative content information rather than task-irrelevant style information. Extensive experiments show that the model generalization ability can be improved significantly in the unseen domain when transferring from SceneFlow to Middlebury, ETH3D and KITTI.
Chihao Ma, Pengcheng Zeng, Jucai Zhai, Yong Zhao 0010, Xinan Wang
ICIP5
2023 MFF: An effective method of solving the ill regions in stereo matching
abstract
Abstract In the current stereo matching field, the accuracy of the derived disparity map is highly dependent on the processing capability of ill regions. Fortunately, we find that the use of local information will eliminate the negative effects associated with ill regions. As a result of the above discovery, we propose the Concatenated Dilated Convolution (CDC) block and the Multi‐scale Feature Fusion module (MFF), which are capable of effectively extracting regional context by increasing the receptive field in parallel and channel‐wise ways. The CDC block can expand the receptive field by applying multiple dilated convolutions at different dilation rates in parallel to enhance the smoothness of the feature map. By constructing parallel CDC blocks in a multiple dilated manner, the MFF module can improve the smoothness of the feature map. In addition, to control the number of parameters in the MFF network, a high‐performance channel‐distribution algorithm is proposed, capable of adjusting the weights of each module and convolution in an adaptive manner while reducing the number of parameters. Extensive experiments have demonstrated that MFF and CDC can effectively improve the performance of ill areas and networks with a minimal number of parameters.
Ren Qian, Renyan Feng, Wangduo Xie, Wenbang Yang, Yong Zhao 0010
IET Comput. Vis.5
2023 A light-weight stereo matching network based on multi-scale features fusion and robust disparity refinement
abstract
Abstract In recent years, convolutional‐neural‐network based stereo matching methods have achieved significant gains compared to conventional methods in terms of both speed and accuracy. Current state‐of‐the‐art disparity estimation algorithms require many parameters and large amounts of computational resources and are not suited for applications on edge devices. In this paper, an end‐to‐end light‐weight network (LWNet) for fast stereo matching is proposed, which consists of an efficient backbone with multi‐scale feature fusion for feature extraction, a 3D U‐Net aggregation architecture for disparity computation, and color guidance in a 2D convolutional neural network (CNN) for disparity refinement. MobileNetV2 is adopted as an efficient backbone in feature extraction. The channel attention module is applied to improve the representational capacity of features and multi‐resolution information is adaptively incorporated into the cost volume via cross‐scale connections. Further, a left‐right consistency check and color guidance refinement are introduced and a robust disparity refinement network is designed with skip connections and dilated convolutions to capture global context information and improve disparity estimation accuracy with little computational cost and memory space. Extensive experiments on Scene Flow, KITTI 2015, and KITTI 2012 demonstrate that the proposed LWNet achieves competitive accuracy and speed when compared with state‐of‐the‐art stereo matching methods.
Yong Zhao 0010, Zhiguo Feng, Haiwei Sang, Zhenbo Zhang, Guiying Zhang
IET Image Process.2
2023 An Algorithm for Single View Occlusion Area Detection in Binocular Stereo Matching
abstract
Occlusion area detection is a crucial step affecting the performance of the binocular stereo matching algorithm, but the traditional method of occlusion area detection has two major problems, including left–right consistency detection (LRC). First, these algorithms must obtain the left and right disparity maps with precision. Second, these algorithms cannot detect the occlusion region at the image’s borders. We propose the single view occlusion area detective (SVOAD) algorithm to detect these occlusion areas and better deal with them. The SVOAD can detect the area of occlusion from a single image, thereby reducing the computational cost. Additionally, the algorithm can detect the occlusion region in all image regions. This paper also improves the guided filter so that it works better with the end-to-end neural network and makes the SVOAD algorithm work better.
Ren Qian, Yong Zhao 0010, Renyan Feng, Wenbang Yang, Zaijun Zhang
Int. J. Pattern Recognit. Artif. Intell.2
2022 EAI-Stereo: Error Aware Iterative Network for Stereo Matching
Haoliang Zhao, Huizhou Zhou, Yongjun Zhang 0007, Yong Zhao 0010, Yitong Yang, Ting Ouyang
ACCV (1)4
2022 OMNET: Real-Time Stereo Matching with Unsupervised Occlusion Mask
abstract
Although CNN has powerful learning capability, it is still difficult for CNN to judge corresponding points in the occlusion region. The ghosting effect in the occlusion region during feature warping is the bottleneck of performance improvement for many stereo matching networks. In this paper, we propose an Occlusion-Aware Refinement Module (OARM), which can learn a rough occlusion map from multi-scale aggregated cost volumes without occlusion supervision to mask and filter pernicious occluded regions in a warped image. Cooperating with well-designed simple yet efficient 2D-based Intra/Cross-Level Aggregation Modules to effectively and efficiently aggregate information of different scales before the disparity refinement stage, OARM helps our proposed OM-Net achieve an error rate of 1.82% (D1-all) on KITTI 2015 dataset, which is even better than most 3D-based networks. Meanwhile, OMNet keeps real-time characteristic and could process a 1248×384 resolution image pair at 26 fps.
Shuiqiang Ye, Xin'an Wang, Yong Zhao 0010
ICIP4
2022 MLP-Stereo: Heterogeneous Feature Fusion in MLP for Stereo Matching
abstract
CNNs’ strong inductive biases of locality and weight sharing provide powerful representation ability and data sample utilization efficiency. However, the weight sharing might smooth out the discrepancy between similar pixels, resulting in the wrong matching between left and right camera-image pair in thin structures region and repetitive texture region. In this paper, we propose a novel Heterogeneous Feature Fusion in MLP (HFF-MLP) for Stereo matching. It employs MLP structure and relaxes the weights sharing in the local spatial region. To this end, pixels in thin structures region and repetitive texture region are dealt with independently using the exclusive weights. Based on HFF -MLP module, we design a real-time network, i.e., MLP-Stereo. Experimental results show that our proposed HFF-MLP achieves competitive results on KITTI 2015 test dataset with the running time of 46 milliseconds. Furthermore, it performs much better than other real-time network in thin structures region and repetitive texture region.
Shuiqiang Ye, Pengcheng Zeng, Xin'an Wang, Yong Zhao 0010
ICIP6
2022 Self-adaptive Multi-scale Aggregation Network for Stereo Matching
abstract
Nowadays stereo matching architectures based on convolutional neural network has achieved remarkable performance. However the existing methods still lack the capability to find the correspondence in ill-posed regions. In this paper we present Self-adaptive Multi-scale Aggregation Network (SMA-Net) for stereo matching. First of all, we construct the cost volume through multi-channels group-wise correlation, which divided the features into groups with different number of channels to enhance the ability of measuring the similarities of features from stereo images. Secondly the self-adaptive cost aggregation is used to regularize the two scale cost volumes from different aggregation branches with intermediate supervise. We conduct comprehensive experiments on SceneFlow, KITTI2012, and KITTI2015 datasets. The competitive results prove that the approach in this paper outperforms many other stereo matching algorithms especially in ill-posed regions.
Shuiqiang Ye, Jiaquan Zhang, Xin'an Wang, Qifei Dai, Zhengzhong Yu, Fuchi Li, Yong Zhao 0010
ICPR8
2022 LESC: Superpixel cut-based local expansion for accurate stereo matching
abstract
Abstract The rapid estimation of the accurate disparity between pixels is the goal of stereo matching. However, it is very difficult for the 3D labels‐based methods due to huge search space of 3D labels, especially for high‐resolution images. In this paper, a novel superpixel cut‐based method is proposed, in an attempt to get the accurate disparity map efficiently, including the multi‐layer superpixel optimization and iteractive local ‐expansion in parallel. As for the multi‐layer superpixel optimization, feature point optimization is designed to get accurate candidate labels that are set for most pixels using non‐local cost aggregation strategy and update per‐pixel labels of the corresponding superpixels from the candidate label sets on the small‐size superpixel layer, and then update the middle to large‐size superpixel layers progressively using non‐local cost aggregation strategy. In order to provide more prior information to identify weak texture and textureless regions in non‐local cost aggregation, the weight combination of “intensity + gradient + binary image” is proposed for constructing an optimal minimum spanning tree (MST) to calculate the aggregated matching cost and obtain the labels of minimum aggregated matching cost. Moreover, the local patch surrounding the corresponding superpixels is designed to accelerate superpixel optimization in parallel, and a neighborhood structure is presented to optimize the algorithm in this study, including superpixel neighborhood and patch neighborhood. As for the iteractive local ‐expansion, three layers of patch structure corresponding to the superpixel neighborhood structure is proposed for optimizing the algorithm in this study. The experimental results show that higher accuracy could be achieved via the method in this study compared with some known state‐of‐the‐art stereo methods on KITTI 2015 and Middlebury benchmark V3, which are the standard benchmarks for testing the stereo matching methods.
Xian Jing Cheng, Yong Zhao 0010, Wenbang Yang, Zhijun Hu, Xiaomin Yu, Haiwei Sang, Guiying Zhang
IET Image Process.2
2022 A novel cell structure-based disparity estimation for unsupervised stereo matching
abstract
Abstract It is well known that preserving depth edges is an effective solution for achieving the accurate disparity map in stereo matching, but many state‐of‐the‐art methods do not preserve depth edges well. In order to solve it efficiently, the cell structure containing irregular and regular shape regions is designed to preserve depth edges. Based on the well‐designed cell structure, a novel disparity estimation method for stereo matching is proposed, in which a two‐layer disparity optimization method is proposed to refine the disparity plane; it includes the front‐parallel disparities computation and slanted‐surfaces disparity plane refinement. In the framework of front‐parallel disparities computation, a tree‐based cost aggregation method is presented to make full use of the segmentation information of cells and then performing semi‐global cost aggregation. In the framework of slanted‐surfaces disparity plane refinement, a new probability model is proposed that employs Bayesian inference for refining disparities in textureless, weak texture and occluded regions. Experimental results show that higher accuracy could be achieved via the proposed method compared with some known state‐of‐the‐art stereo methods on KITTI 2015 and Middlebury dataset, which are the standard benchmarks for testing the stereo matching methods. It can also be indicated that the proposed method can produce accurate disparity map and have good generalization performance.
Xian Jing Cheng, Yong Zhao 0010, Wenbang Yang, Zhijun Hu, Xiaomin Yu, Haoliang Zhao, Pengcheng Zeng
IET Image Process.2
2022 Edge supervision and multi-scale cost volume for stereo matching
Zhiguo Feng, Yong Zhao 0010, Guiying Zhang
Image Vis. Comput.3
2021 A Unified Multi-Scenario Attacking Network for Visual Object Tracking
abstract
Existing methods of adversarial attacks successfully generate adversarial examples to confuse Deep Neural Networks (DNNs) of image classification and object detection, resulting in wrong predictions. However, these methods are difficult to attack models of video object tracking, because the tracking algorithms could handle sequential information across video frames and the categories of targets tracked are normally unknown in advance. In this paper, we propose a Unified and Effective Network, named UEN, to attack visual object tracking models. There are several appealing characteristics of UEN: (1) UEN could produce various invisible adversarial perturbations according to different attack settings by using only one simple end-to-end network with three ingenious loss function; (2) UEN could generate general visible adversarial patch patterns to attack the advanced trackers in the real-world; (3) Extensive experiments show that UEN is able to attack many state-of-the-art trackers effectively (e.g. SiamRPN-based networks and DiMP) on popular tracking datasets including OTB100, UAV123, and GOT10K, making online real-time attacks possible. The attack results outperform the introduced baseline in terms of attacking ability and attacking efficiency.
Xuesong Chen 0001, Canmiao Fu, Feng Zheng 0001, Yong Zhao 0010, Hongsheng Li 0001, Ping Luo 0002, Guo-Jun Qi
AAAI4
2021 HPA-Net: Hierarchical and Parallel Aggregation Network for Context Learning in Stereo Matching
Yong Zhao 0010
CAIP (1)4
2021 Hierarchical Context Guided Aggregation Network for Stereo Matching
abstract
Nowadays, CNN-based stereo matching methods achieved remarkable performance, and how to efficiently exploit contextual information in cost aggregation stage is the key to improve performance. In this paper, we propose a simple yet efficient network named Hierarchical Context Guided Aggregation Network (HCGANet). Specifically, a novel cost aggregation module is developed to replace widely used 3D convolutions. Firstly, we construct pyramid cost volumes which carry multi-level distinctive and discriminative representation. Additionally, an intra-level aggregation module is presented for single-level regularization and contextual information learning. Moreover, we develop an inter-level aggregation module to hierarchically regularize cost volumes via the guidance from coarser scales. The proposed aggregation module is lightweight and complementary, further improving the robustness and performance of disparity estimation. Extensive experiments demonstrate that the proposed method achieves superior results for both efficiency and accuracy on Scene-Flow and KITTI benchmarks.
Wangduo Xie, Zijing Huang, Yong Zhao 0010
ICASSP5
2021 Hierarchical and Multi-Level Cost Aggregation For Stereo Matching
abstract
Nowadays, convolutional neural networks based on deep learning have greatly improved the performance of stereo matching. To obtain higher disparity estimation accuracy in ill-posed regions, this paper proposes a hierarchical and multi-level model based on a novel cost aggregation module (HMLNet). This effective cost aggregation consists of two main modules: one is the multi-level cost aggregation which incorporates global context information by fusing information in different levels, and the other called the hourglass+ module utilizes sufficiently volumes in the same level to regularize cost volumes better. Also, we take advantage of disparity refinement with residual learning to boost robustness to challenging situations. We conducted comprehensive experiments on Sceneflow, KITTI 2012, and KITTI 2015 datasets. The competitive results prove that our approach outperforms many other stereo matching algorithms.
Fukun Xia, Yong Zhao 0010
ICIP5
2021 MPANET: Multi-Scale Pyramid Aggregation Network For Stereo Matching
abstract
At present, the performance of the end-to-end stereo matching networks based on CNN greatly exceed the traditional stereo matching networks, but the accuracy in those ill-posed regions like foreground areas is still not optimistic. In this paper, we propose a novel design to improve the prediction performance of disparity in foreground. First, a multi-scale pyramid aggregation module with hourglass-like structure is designed to effectively utilize the aggregation information of different scales. In addition, we propose MPANet, a novel end-to-end stereo matching network, which significantly alleviates the disparity mismatch problem in foreground areas. Experimental results on Scene Flow and KITTI datasets also demonstrate that our network has competitive performance among existing first-class methods.
Yong Zhao 0010
ICIP5
2021 Fast Multi-Scale Residual Fusion Network for Stereo Matching
abstract
Recent deep convolution-based stereo matching methods have shown significant progress. However, most state-of-the-art models achieve high accuracy by using 3D convolutions during cost aggregation, which come with more floating-point computations and makes it difficult for deployment in real-time applications. In this paper, we propose an effective and efficient multi-scale aggregation module (without 3D convolution) to build our fast Multi-Scale Residual Fusion Network (MSRFNet). Three lightweight components are involved in our aggregation module: combination block, attention block, and residual fusion block. The combination block combines features with different receptive fields and the attention block emphasizes salient regions in various cost volumes. The residual fusion module focuses on extracting the differences between adjacent cost volumes, using progressively residual aggregation instead of simply stacking or adding. Extensive experiments on Scene Flow and KITTI benchmarks demonstrate that our method achieves competitive accuracy com-pared with state-of-the-art methods while running at 56ms.
Zijing Huang, Wangduo Xie, Yong Zhao 0010
ICME5
2021 Intragroup sparsity for efficient inference
abstract
This work studies intragroup sparsity, a fine-grained structural constraint on network weight parameters. It eliminates the computational inefficiency of fine-grained sparsity due to irregular dataflow, while at the same time achieving high inference accuracy. We present theoretical analysis on how weight group sizes affect sparsification error, and on how the performance of pruned networks changes with sparsity level. Further, we analyze inference-time I/O cost of two different strategies for achieving intragroup sparsity and how the choice of strategies affect I/O cost under mild assumptions on accelerator architecture. Moreover, we present a novel training algorithm that yield models of improved accuracies over the standard training approach under the intragroup sparsity constraint.
Zilin Yu, Chao Wang 0037, Yong Zhao 0010, Xundong Wu
IJCNN4
2020 Salience-Guided Cascaded Suppression Network for Person Re-Identification
abstract
Employing attention mechanisms to model both global and local features as a final pedestrian representation has become a trend for person re-identification (Re-ID) algorithms. A potential limitation of these methods is that they focus on the most salient features, but the re-identification of a person may rely on diverse clues masked by the most salient features in different situations, e.g., body, clothes or even shoes. To handle this limitation, we propose a novel Salience-guided Cascaded Suppression Network (SCSN) which enables the model to mine diverse salient features and integrate these features into the final representation by a cascaded manner. Our work makes the following contributions: (i) We observe that the previously learned salient features may hinder the network from learning other important information. To tackle this limitation, we introduce a cascaded suppression strategy, which enables the network to mine diverse potential useful features that be masked by the other salient features stage-by-stage and each stage integrates different feature embedding for the last discriminative pedestrian representation. (ii) We propose a Salient Feature Extraction (SFE) unit, which can suppress the salient features learned in the previous cascaded stage and then adaptively extracts other potential salient feature to obtain different clues of pedestrians. (iii) We develop an efficient feature aggregation strategy that fully increases the network’s capacity for all potential salience features. Finally, experimental results demonstrate that our proposed method outperforms the state-of-the-art methods on four large-scale datasets. Especially, our approach exceeds the current best method by over 7% on the CUHK03 dataset.
Xuesong Chen 0001, Canmiao Fu, Yong Zhao 0010, Feng Zheng 0001, Jingkuan Song, Rongrong Ji, Yi Yang 0001
CVPR3
2020 One-Shot Adversarial Attacks on Visual Tracking With Dual Attention
abstract
Almost all adversarial attacks in computer vision are aimed at pre-known object categories, which could be offline trained for generating perturbations. But as for visual object tracking, the tracked target categories are normally unknown in advance. However, the tracking algorithms also have potential risks of being attacked, which could be maliciously used to fool the surveillance systems. Meanwhile, it is still a challenging task that adversarial attacks on tracking since it has the free-model tracked target. Therefore, to help draw more attention to the potential risks, we study adversarial attacks on tracking algorithms. In this paper, we propose a novel one-shot adversarial attack method to generate adversarial examples for free-model single object tracking, where merely adding slight perturbations on the target patch in the initial frame causes state-of-the-art trackers to lose the target in subsequent frames. Specifically, the optimization objective of the proposed attack consists of two components and leverages the dual attention mechanisms. The first component adopts a targeted attack strategy by optimizing the batch confidence loss with confidence attention while the second one applies a general perturbation strategy by optimizing the feature loss with channel attention. Experimental results show that our approach can significantly lower the accuracy of the most advanced Siamese network-based trackers on three benchmarks.
Xuesong Chen 0001, Xiyu Yan, Feng Zheng 0001, Yong Jiang 0001, Shutao Xia, Yong Zhao 0010, Rongrong Ji
CVPR6
2020 Dedge-AGMNet: An Effective Stereo Matching Network Optimized by Depth Edge Auxiliary Task
abstract
To improve the performance in ill-posed regions, this paper proposes an atrous granular multi-scale network based on depth edge subnetwork(Dedge-AGMNet). According to a general fact, the depth edge is the binary semantic edge of instance-sensitive. This paper innovatively generates the depth edge ground-truth by mining the semantic and instance dataset simultaneously. To incorporate the depth edge cues efficiently, our network employs the hard parameter sharing mechanism for the stereo matching branch and depth edge branch. The network modifies SPP to Dedge-SPP, which fuses the depth edge features to the disparity estimation network. The granular convolution is extracted and extends to 3D architecture. Then we design the AGM module to build a more suitable structure. This module could capture the multi-scale receptive field with fewer parameters. Integrating the ranks of different stereo datasets, our network outperforms other stereo matching networks and advances state-of-the-art performances on the Sceneflow, KITTI 2012 and KITTI 2015 benchmark datasets.
Weida Yang, Xindong Ai, Zuliu Yang, Yong Xu 0001, Yong Zhao 0010
ECAI5
2020 Non-Local Nested Residual Attention Network for Stereo Image Super-Resolution
abstract
Nowadays CNN-based stereo image super-resolution(SR) methods have obtained remarkable performance. However, most of existing methods only superficially portrayed the low layer features without considering the uneven distribution of information, which is insufficient because stereo image warping and sub-pixel upsampling require discriminative features to identify corresponding pixels. To address this problem, in this paper, we propose a novel network named Non-local Nested Residual Attention Network (NNRANet). Specifically, a non-local dilated attention module (NDAM) is developed to exploit the rich hierarchical feature and capture the long-range dependencies between pixels. Moreover, we present a nested residual group (NRG) with dense connections and multiple nested residual sub-network, which not only continuously remembers and extracts the stereo fusion feature, but also enables training a deeper and more stable network. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on the Middlebury, KITTI 2012 and KITTI 2015 datasets.
Wangduo Xie, Zhisheng Lu, Yong Zhao 0010
ICASSP5
2020 Hijacking Tracker: A Powerful Adversarial Attack on Visual Tracking
abstract
Visual object tracking has made important breakthroughs with the assistance of deep learning models. Unfortunately, recent research has clearly proved that deep learning models are vulnerable to malicious adversarial attacks, which mislead the models making wrong decisions by perturbing the input image. The threat to the models alerts us to pay attention to the model security of deep learning- based tracking algorithms. Therefore, we study the adversarial attacks against advanced trackers based on deep learning to better identify the vulnerability of tracking algorithms. In this paper, we propose to add slight adversarial perturbations to the input image by an inconspicuous but powerful attack strategy-hijacking algorithm. Specifically, the hijacking strategy misleads trackers in two aspects: one is shape hijacking that changes the shape of the model output; the other is position hijacking that gradually pushes the output to any position in the image frame. Besides, we further propose an adaptive optimization approach to integrate two hijacking mechanisms efficiently. Eventually, the hijacking algorithm results in fooling the tracker to track the wrong target gradually. The experimental results demonstrate the powerful attack ability of our method-quickly hijacking state-of-the-art trackers and reducing the accuracy of these models by more than 90% on OTB2015.
Xiyu Yan, Xuesong Chen 0001, Yong Jiang 0001, Shutao Xia, Yong Zhao 0010, Feng Zheng 0001
ICASSP5
2020 Suppressing Features That Contain Disparity Edge For Stereo Matching
abstract
Existing networks for stereo matching usually use 2-D CNN as the feature extractor. However, objects are usually continuous in spatial, if an extracted feature contains disparity edge (the representation of this feature on original image contains disparity edge), then this feature usually not occur inside the region of an object. We propose a novel attention mechanism to suppress features containing disparity edge, named SDE-Attention (SDEA). We notice that features containing disparity edge are usually continuous in one image and discontinuous in another, which means that they usually have a greater difference in two feature maps of the same layer than features that don't contain disparity edge. SDEA calculate the weight matrix of the intermediate feature map according to this trait, then the weight matrix is multiplied to the intermediate feature map. We test SDEA on PSMNet, experimental results show that our method has a significant improvement in accuracy and our network achieves state-of-the-art performance among the published networks.
Xindong Ai, Zuliu Yang, Weida Yang, Yong Zhao 0010, Zhengzhong Yu, Fuchi Li
ICPR4
2020 Deeply-fused Attentive Network for Stereo Matching
abstract
In this paper, we propose a novel learning-based network for stereo matching called DF-Net, which makes three main contributions that are experimentally shown to have practical merit. Firstly, we further increase the accuracy by using the deeply fused spatial pyramid pooling (DF-SPP) module, which can acquire the continuous multi-scale context information in both parallel and cascade manners. Secondly, we introduce channel attention block to dynamically boost the informative features. Finally, we propose a stacked encoder-decoder structure with 3D attention gate for cost regularization. More precisely, the module fuses the coding features to their next encoder-decoder structure under the supervision of attention gate with long-range skip connection, and thus exploit deep and hierarchical context information for disparity prediction. The performance on SceneFlow and KITTI datasets shows that our model is able to generate better results against several state-of-the-art algorithms.
Zuliu Yang, Xindong Ai, Weida Yang, Yong Zhao 0010, Qifei Dai, Fuchi Li
ICPR4
2020 MPA-Net: multi-path attention stereo matching network
abstract
A novel learning‐based end‐to‐end network for stereo matching, named Multi‐path Attention Stereo Matching (MPA‐Net), is introduced in this study. Different from existing methods, the multi‐path attention aggregation module is designed firstly, named MPA, which is a unified structure using three different parallel layers with a respective attention mechanism to extract the multi‐scale informational features. Secondly, the method of cost volume construction, which differs from the traditional stereo matching methods, is extended. And then, the absolute difference between two input features is calculated. Furthermore, a u‐shaped structure with 3D attention gate is selected as the encoder‐decoder module. Specifically, the module is used to fuse the encoding features to their corresponding decoding features under the supervision of the authors' attention gate with skip‐connection, and thus exploit more significant information for matching cost regularisation and disparity prediction. Finally, specific experiments are conducted to evaluate their network on SceneFlow, KITTI2012 and KITTI2015 data sets. The results show that their method achieves a better improvement in disparity maps prediction compared with some existing state‐of‐the‐art methods on KITTI benchmark.
Haiwei Sang, Zuliu Yang, Yong Zhao 0010
IET Image Process.4
2020 PCANet: Pyramid convolutional attention network for semantic segmentation
abstract
Pyramid Convolutional Attention Network is proposed to efficiently capture long-range dependency and fuse features from different levels for benefitting semantic segmentation problems. In this paper, we focus on how to extract more representative features for segmentation object recognition and design a decoder to recover details in a more efficient way. Inspired by atrous sampling and attention mechanism, we propose Pyramid Atrous Attention module to capture long-range dependency for learning richer contextual features. We also find that features of different levels have diverse representation so we design Convolutional Attention Refinement module to provide global context for low-level features and local details for high-level features. By combining with these two efficient module, we construct our Pyramid Convolutional Attention Network (PCANet), which achieves state-of-the-art results on Pascal VOC 2012 and Cityscapes benchmark.
Haiwei Sang, Qiuhao Zhou, Yong Zhao 0010
Image Vis. Comput.3
2020 Improving the Accuracy of Binarized Neural Networks and Application on Remote Sensing Data
abstract
Deep neural networks are well known to achieve outstanding results in many domains. Recently, many researchers have introduced deep neural networks into remote sensing (RS) data processing. However, typical RS data usually possesses enormous scale. Processing RS data with deep neural networks requires a rather demanding computing hardware. Most high-performance deep neural networks are associated with highly complex network structures with many parameters. This restricts their deployment for real-time processing in satellites. Many researchers have attempted overcoming this obstacle by reducing network complexity. One of the promising approaches able to reduce network computational complexity and memory usage dramatically is network binarization. In this letter, through analyzing the learning behavior of binarized neural networks (BNNs), we propose several novel strategies for improving the performance of BNNs. Empirical experiments prove these strategies to be effective in improving BNN performance for image classification tasks on both small- and large-scale data sets. We also test BNN on a remote sense data set with positive results. A detailed discussion and preliminary analysis of the strategies used in the training are provided.
Qing Wu 0008, Congchong Chen, Chao Wang 0037, Yong Zhao 0010, Xundong Wu
IEEE Geosci. Remote. Sens. Lett.5
2020 Moving Cast Shadows Segmentation Using Illumination Invariant Feature
abstract
This paper presents an effective framework for removing moving cast shadows. Taking the reflection property of object surface for shadow regions under static and fixed scenes, an approximation estimation strategy of bidirectional reflectance distribution function as illumination invariant feature is proposed. It is valid for different types of shadow scenes. In this paper, we propose a new multiple ratios-based technique to justify shadow type for each frame: intensity ratio, area ratio and edge ratio of shadow regions are introduced. According to shadow types, several specified strategies are designed. For weak shadows, multiple features fusion strategy is employed, including color constancy, texture consistency and illumination invariant. For strong shadows, illumination invariant is utilized to detect the umbra and color constancy is utilized to detect the penumbra. Moreover, a suite of shadow direction features is firstly proposed to identify penumbra. The proposed approach is verified in fourteen video sequences varying from weak to strong shadows. The experimental results demonstrate the effectiveness and robustness of the proposed method for both indoor and outdoor scenes compared with some state-of-the-art approaches.
Bingshu Wang, Yong Zhao 0010, C. L. Philip Chen
IEEE Trans. Multim.2
2019 Scanet: Spatial-channel Attention Network for 3D Object Detection
abstract
This paper aims to achieve high-accuracy 3D object detection, in which we propose a novel Spatial-Channel Attention Network (SCANet), a two-stage detector that takes both LIDAR point clouds and RGB images as input to generate 3D object estimates. The first stage is a 3D region proposal network (RPN) in which we put forward a new Spatial-Channel Attention (SCA) module and an Extension Spatial Upsample (ESU) module. Using the pyramid pooling structure and global average pooling, the SCA module can not only effectively incorporate multi-scale and global context information, but also produce spatial and channel-wise attention to select discriminative features. The ESU module in the decoder can recover the lost spatial information caused by consecutive pooling operators to generate reliable 3D region proposals. In the second stage, we design a new multi-level fusion scheme for accurate classification and 3D bounding box regression. Experimental results demonstrate that SCANet achieves state-of-the-art performance on the challenging KITTI 3D object detection benchmark.
Haihua Lu, Xuesong Chen 0001, Guiying Zhang, Qiuhao Zhou, Yanbo Ma, Yong Zhao 0010
ICASSP6
2019 Multi-attention Network for Thoracic Disease Classification and Localization
abstract
The chest X-ray is one of the most commonly available radiological examinations for diagnosing lung diseases. This task remains a major challenge due to 1) the shortage of accurate annotations for chest X-ray examinations, 2) the diversity of lesion areas on X-rays from different thoracic disease and 3) the problem of class imbalance in existing chest X-ray databases. In this paper, we propose a new multi-attention convolutional neural network for thoracic disease classification and localization. First, the framework is equipped with squeeze-and-excitation (SE) block as a feature attention module to offer a chance of cross-channel feature recalibration. Second, we propose a novel space attention module to combine global and local information. Third, we present a hard examples attention module to alleviate the class imbalance problem. The comprehensive experiments are performed on the ChestX-ray14 dataset. Quantitative and qualitative results demonstrate that our method outperforms the state-of-the-art algorithm.
Yanbo Ma, Qiuhao Zhou, Xuesong Chen 0001, Haihua Lu, Yong Zhao 0010
ICASSP5
2019 Discriminative Features Reconstruction Network for Semantic Segmentation
abstract
Thanks to the development of convolutional neural networks (CNNs), researchers have proposed lots of effective semantic segmentation models. However, there are still two problems disturbing researchers, one of which is objects misidentification on the image level and another one is poor performance on details, especially the boundary of objects. To tackle these two problems, we propose Discriminative Features Reconstruction Network (DFR) containing two modules: Second-order Pyramid Features Reconstruction Module (SPFR) and Second-order Boundary Attention Module (SBA). Specifically, SPFR fuses different scales features to gain pyramid receptive field. Besides, SPFR extracts second-order statistics data to retrieve more discriminative features. Furthermore, we put forward SBA that is helpful to refine the segmentation results. On SBA, low-level features recover localization details under the high-level feature guidance. Our DFR achieves state-of-the-art performance on PASCAL VOC 2012 dataset with mIoU accuracy 81.1% without pre-training COCO dataset and post-processing.
Qiuhao Zhou, Yanbo Ma, Haihua Lu, Xuesong Chen 0001, Yong Zhao 0010
ICASSP5
2019 Non-Local Recurrent Neural Memory for Supervised Sequence Modeling
abstract
Typical methods for supervised sequence modeling are built upon the recurrent neural networks to capture temporal dependencies. One potential limitation of these methods is that they only model explicitly information interactions between adjacent time steps in a sequence, hence the high-order interactions between nonadjacent time steps are not fully exploited. It greatly limits the capability of modeling the long-range temporal dependencies since one-order interactions cannot be maintained for a long term due to information dilution and gradient vanishing. To tackle this limitation, we propose the Non-local Recurrent Neural Memory (NRNM) for supervised sequence modeling, which performs non-local operations to learn full-order interactions within a sliding temporal block and models the global interactions between blocks in a gated recurrent manner. Consequently, our model is able to capture the long-range dependencies. Besides, the latent high-level features contained in high-order interactions can be distilled by our model. We demonstrate the merits of our NRNM approach on two different tasks: action recognition and sentiment analysis.
Canmiao Fu, Wenjie Pei, Qiong Cao, Chaopeng Zhang, Yong Zhao 0010, Xiaoyong Shen, Yu-Wing Tai
ICCV5
2019 Sequentially Refined Spatial and Channel-Wise Feature Aggregation in Encoder-Decoder Network for Single Image Dehazing
abstract
Single image haze removal is a challenging problem due to its inherent ill-posed nature. Several prior-based and learning-based methods have been proposed to solve this problem and they have achieved superior results. However, the performance of prior-based methods is limited by hand-designed features. Meanwhile, the learning-based methods have flaws in spatial correlation. Because they just employ the convolutional neural network (CNN) as an end-to-end mapping module, generating dehazed image directly or learning parameters in atmospheric scattering model, without associating the relationship between feature activation and distribution of haze. Instead of using prior or CNN to simply estimate parameters of atmospheric scattering model, we propose a sequentially refined Spatial and Channel-wise Feature Aggregation (SCFA) dehazing network, called SCFADN. Specifically, our network learns residues between clean images and hazy images through an improved encoder-decoder network which incorporated the proposed sequentially refined SCFA module. This module efficiently learns long-range feature correlation for dehazing residue modeling, which enables the shallow network removing haze effectively. Extensive experiments on synthetic and real datasets results demonstrate that the proposed method achieves superior performance over the state-of-the-art methods.
Xuesong Chen 0001, Haihua Lu, Kaili Cheng, Yanbo Ma, Qiuhao Zhou, Yong Zhao 0010
ICIP6
2019 ACNet: Aggregated Channels Network for Automated Mitosis Detection
Kaili Cheng, Xuesong Chen 0001, Yanbo Ma, Mengjie Bai, Yong Zhao 0010
PAKDD (1)6
2019 Pedestrian detection using multi-channel visual feature fusion by learning deep quality model
Peijia Yu, Yong Zhao 0010, Jing Zhang 0022, Xiaoyao Xie
J. Vis. Commun. Image Represent.2
2018 CHS-NET: A Cascaded Neural Network with Semi-Focal Loss for Mitosis Detection
abstract
Counting of mitotic figures in hematoxylin and eosin(H&E) stained histological slide is the main indicator of tumor proliferation speed which is an important biomarker indicative of breast cancer patients’ prognosis. It is difficult to detect mitotic cells due to the diversity of the cells and the problem of class imbalance. We propose a new network called CHS-NET which is a cascaded neural network with hard example mining and semi-focal loss to detect mitotic cells in breast cancer. First, we propose a screening network to identify the candidates of mitotic cells preliminary and a refined network to identify mitotic cells from these candidates more accurately. We propose a new feature fusion module in each network to explore complex nonlinear predictors and improve accuracy. Then, we propose a novel loss named semi-focal loss and we use off-line hard example mining to solve the problem of class imbalance and error labeling. Finally, we propose a new training skill of cutting patches in the whole slide image, considering the size and distribution of mitotic cells. Our method achieves 0.68 F1 score which outperforms the best result in Tumor Proliferation Assessment Challenge 2016 held by MICCAI.
Yanbo Ma, Qiuhao Zhou, Kaili Cheng, Xuesong Chen 0001, Yong Zhao 0010
ACML6
2018 Depth Super-Resolution Using Joint Adaptive Weighted Least Squares And Patching Gradient
abstract
This paper presents a flexible framework for the challenging task of color-guided depth upsampling. Some state-of-the-art approaches apply an aligned RGB image for depth recovery. Unfortunately, these kinds of methods may result in texture copying artifacts and edge blurring artifacts. To address these difficulties, we propose an adaptive weighted least squares framework of choosing different guidance weight for variant conditions flexibly. First of all, in the framework, we propose a joint adaptive color weighting scheme in which the depth maps and color images jointly choose a proper weight term for diverse cases. Then, a patch-based smoothness measuring approach called patching-gradient method (PGM) is proposed to distinguish the discontinuities and smooth areas. Our PGM is robust to dense noise and preserve weak edges effectively. Quantitative and qualitative experiments on noisy T of -like datasets demonstrate our frameworks effectiveness on suppressing both texture copying artifacts and edge blurring artifacts.
Bingshu Wang, Yong Zhao 0010
ICASSP4
2018 Hard Shadows Removal Using an Approximate Illumination Invariant
abstract
Hard shadows detection and removal from foreground masks is a challenging step in change detection. This paper gives a simple and effective method to address hard shadows. There are inside portion and boundary portion in hard shadows. Pixel-wise neighborhood ratio is calculated to remove the most of inside shadow points. For the boundaries of shadow regions, we take advantage of color constancy to eliminate the edges of hard shadows and obtain relative accurate objects contours. Then, morphology processing is explored to enhance the integrity of objects. The main contribution of this paper is to design an approximate estimation strategy for illumination invariant based on Lambertian reflectance model without prior knowledge. The proposed method is unsupervised and experimental results on six challenging sequences show the effectiveness and robustness of our approach.
Bingshu Wang, C. L. Philip Chen, Yong Zhao 0010
ICASSP4
2018 Singular value decomposition based virtual representation for face recognition
Guiying Zhang, Wenbin Zou, Xianjie Zhang, Yong Zhao 0010
Multim. Tools Appl.4
2017 A Robust and Efficient Approach to License Plate Detection
abstract
This paper presents a robust and efficient method for license plate detection with the purpose of accurately localizing vehicle license plates from complex scenes in real time. A simple yet effective image downscaling method is first proposed to substantially accelerate license plate localization without sacrificing detection performance compared with that achieved using the original image. Furthermore, a novel line density filter approach is proposed to extract candidate regions, thereby significantly reducing the area to be analyzed for license plate localization. Moreover, a cascaded license plate classifier based on linear support vector machines using color saliency features is introduced to identify the true license plate from among the candidate regions. For performance evaluation, a data set consisting of 3977 images captured from diverse scenes under different conditions is also presented. Extensive experiments on the widely used Caltech license plate data set and our newly introduced data set demonstrate that the proposed approach substantially outperforms state-of-the-art methods in terms of both detection accuracy and run-time efficiency, increasing the detection ratio from 91.09% to 96.62% while decreasing the run time from 672 to 42 ms for processing an image with a resolution of 1082×728 . The executable code and our collected data set are publicly available.
Yule Yuan, Wenbin Zou, Yong Zhao 0010, Xin'an Wang, Xuefeng Hu, Nikos Komodakis
IEEE Trans. Image Process.3
2017 An Adaptive Background Modeling Method for Foreground Segmentation
abstract
Background modeling has played an important role in detecting the foreground for video analysis. In this paper, we presented a novel background modeling method for foreground segmentation. The innovations of the proposed method lie in the joint usage of the pixel-based adaptive segmentation method and the background updating strategy, which is performed in both pixel and object levels. Current pixel-based adaptive segmentation method only updates the background at the pixel level and does not take into account the physical changes of the object, which may result in a series of problems in foreground detection, e.g., a static or low-speed object is updated too fast or merely a partial foreground region is properly detected. To avoid these deficiencies, we used a counter to place the foreground pixels into two categories (illumination and object). The proposed method extracted a correct foreground object by controlling the updating time of the pixels belonging to an object or an illumination region respectively. Extensive experiments showed that our method is more competitive than the state-of-the-art foreground detection methods, particularly in the intermittent object motion scenario. Moreover, we also analyzed the efficiency of our method in different situations to show that the proposed method is available for real-time applications.
Zuofeng Zhong, Bob Zhang 0001, Guangming Lu 0002, Yong Zhao 0010, Yong Xu 0001
IEEE Trans. Intell. Transp. Syst.4
2016 Background subtraction using dual-class backgrounds
abstract
This paper presents a novel approach to background subtraction which aims to extract moving objects in video stream. To this end, a novel background model is proposed by using both working backgrounds and candidate backgrounds, which can be transferred to each other according to an adaptive mechanism. The input image (video frame) is compared and evaluated with these dual-class backgrounds (DCB) to detect foreground objects. Furthermore, for robust background modeling a novel background updating scheme is proposed based on the life-value which represents the existing time of a background sample, and the access-time which represents the number of valid visits of a background sample. Experiments on a standard dataset demonstrated the effectiveness and robustness of the proposed approach by comparing it with the previous typical background subtraction techniques.
Bingshu Wang, Wenqian Zhu, Yong Zhao 0010, Wenbin Zou
ICARCV4
2015 Robust and Real-Time Lane Marking Detection for Embedded System
Yueting Guo, Yongjun Zhang 0007, Yong Zhao 0010
ICIG (3)5
2015 A Vision-Based Method for Parking Space Surveillance and Parking Lot Management
Yangyang Hu, Xuefeng Hu, Yong Zhao 0010
ICIG (1)4
2015 A Foreground Extraction Method by Using Multi-Resolution Edge Aggregation Algorithm
Wenqian Zhu, Bingshu Wang, Xuefeng Hu, Yong Zhao 0010
ICIG (1)4
2015 Unsupervised Joint Salient Region Detection and Object Segmentation
abstract
This paper presents a novel unsupervised algorithm to detect salient regions and to segment out foreground objects from background. In contrast to previous unidirectional saliency-based object segmentation methods, in which only the detected saliency map is used to guide the object segmentation, our algorithm mutually exploits detection/segmentation cues from each other. To achieve this goal, an initial saliency map is generated by the proposed segmentation driven low-rank matrix recovery model. Such a saliency map is exploited to initialize object segmentation model, which is formulated as energy minimization of Markov random field. Mutually, the quality of saliency map is further improved by the segmentation result, and serves as a new guidance for the object segmentation. The optimal saliency map and the final segmentation are achieved by jointly optimizing the defined objective functions. Extensive evaluations on MSRA-B and PASCAL-1500 datasets demonstrate that the proposed algorithm achieves the state-of-the-art performance for both the salient region detection and the object segmentation.
Wenbin Zou, Zhi Liu 0003, Kidiyo Kpalma, Joseph Ronsin, Yong Zhao 0010, Nikos Komodakis
IEEE Trans. Image Process.5
2012 Theory and verification of operator design methodology
Ziyi Hu, Yong Zhao 0010, Xin'an Wang, Ru Huang 0001, Xing Zhang 0002
Sci. China Inf. Sci.2
2011 Real-Time Multiple Vehicles Tracking with Occlusion Handling
abstract
Vehicle detection and tracking is fundamental to vision-based traffic applications. We introduce a real-time system for multiple vehicles tracking with occlusion handling. Firstly, a method of three-level noise removing and vehicle segmentation is presented. Then the segmented vehicles are tracked by using an approach based on normalized area of intersection and the system keeps tracking the occluded vehicles independently by local corner features matching and tracking. Experimental results from video sequences of real-world traffic scenes in the daytime are presented which demonstrate the effectiveness of the algorithm.
Yong Zhao 0010, Yule Yuan
ICIG2