Zhengfa Liang

dblp:121/9474 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0001-7514-4250ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Gated Cross-Attention Network for Depth Completion
abstract
Depth completion is a popular research direction in the field of depth estimation. The fusion of color and depth features is the critical challenge in this task, mainly due to the asymmetry between the rich scene details in color images and the sparse pixels in depth maps. To tackle this issue, we design an efficient Gated Cross-Attention Network that propagates confidence via a gating mechanism, simultaneously extracting and refining key information in both color and depth branches to achieve local spatial feature fusion. Additionally, we incorporate a Transformer-based attention network in low-dimensional space to effectively fuse global features and increase the network’s receptive field. At the same time, we use the Ray Tune mechanism with the AsyncHyperBandScheduler and the HyperOptSearch algorithm to automatically search for the optimal number of module iterations, which also allows us to achieve performance comparable to state-of-the-art methods. We conduct experiments on both indoor and outdoor scene datasets. Our fast network ranked first among real-time methods (below 30ms and 100ms), and our accurate network ranked first among all methods on the KITTI official website at the time of submission.
Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang
ICASSP6
2025 Hierarchical Neural Architecture Search for Fast and Accurate Depth Completion
Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang, Yu-Lin He
ICMR6
2024 Don't Turn a Blind Eye to Localization Noise: Localization Pseudo-label Correction and Learning for Semi-Supervised Object Detection
abstract
Pseudo-labeling has proven to be a simple yet effective technique for semi-supervised object detection (SSOD). However, the inevitable noise problem in pseudo-labels seriously hinders SSOD methods. Existing methods primarily focus on classification noise, while the specific and non-negligible localization noise remains not well-addressed. This paper analyzes the localization noise arising from the alternating learning and generation phases. For the generation phase, we innovatively explore the self-correction ability of models, stepping beyond the simple pseudo-label selection. We propose a localization pseudo-label correction (LPC) strategy to self-correct pseudo boxes and enhance prediction stability. In the learning phase, we propose a noisy localization loss (NLL) to enlarge the penalty of inconsistent predictions, thereby improving localization accuracy. Applied to two classic SSOD methods (Soft Teacher and Unbiased Teacher) and a recent state-of-the-art method (PseCo), our approach consistently improves accuracy across all of them.
Yu-Lin He, Wei Chen 0009, Zhengfa Liang, Ke Liang 0006, Yusong Tan, Yulan Guo
ICME3
2024 Crots: Cross-Domain Teacher-Student Learning for Source-Free Domain Adaptive Semantic Segmentation
Xin Luo 0009, Wei Chen 0009, Zhengfa Liang, Chen Li 0034
Int. J. Comput. Vis.3
2023 Adaptive Scale and Spatial Aggregation for Real-Time Object Detection
abstract
Cutting-edge real-time detectors usually reach real-time performance by adopting lightweight architectures. The accuracy of detection may be limited by their insufficient capabilities to obtain powerful feature representation, which is a notoriously onerous task in machine vision applications. Aiming at this problem, this study proposes a method of adaptive aggregation of features at both scale and spatial levels in an anchor-free framework: 1) at the scale level, a Multi-scale Point Feature Fusion (MPFF) module has been proposed to fuse point features from multiple scales via a self-adaptive re-weighting manner; 2) at the spatial level, a Restrained Deformable Convolution (R-DCN) has been designed to focus on the most informative features in a pre-defined region while avoiding the remote feature distraction. Based on R-DCN, an Adaptive Spatial Aggregation (ASA) module has been presented to alleviate the feature misalignment problem in classification and regression tasks via their respective spatial divisions. Extensive experimental results on MS COCO indicate that Adaptive Aggregation Detector (AADet) achieves a state-of-the-art detection performance, i.e., 41.8 AP at 60 FPS.
Wei Chen 0009, Yu-Lin He, Zhengfa Liang, Yulan Guo
ICASSP3
2023 Adversarial style discrepancy minimization for unsupervised domain adaptation
abstract
Mainstream unsupervised domain adaptation (UDA) methods align feature distributions across different domains via adversarial learning. However, most of them focus on global distribution alignment, ignoring the fine-grained domain discrepancy. Besides, they generally require auxiliary models, bringing extra computation costs. To tackle these issues, this study proposes an UDA method that differentiates individual samples without the help of extra models. To this end, we introduce a novel discrepancy metric, termed style discrepancy, to distinguish different target samples. We also propose a paradigm for adversarial style discrepancy minimization (ASDM). Specifically, we fix the parameters of the feature extractor and maximize style discrepancy to update the classifier, which helps detect more hard samples. Adversely, we fix the parameters of the classifier and minimize the style discrepancy to update the feature extractor, pushing those hard samples near the support of the source distribution. Such adversary helps to progressively detect and adapt more hard samples, leading to fine-grained domain adaptation. Experiments on different UDA tasks validate the effectiveness of ASDM. Overall, without any extra models, ASDM reaches a 46.9% mIoU in the GTA5 to Cityscapes benchmark and an 84.7% accuracy in the VisDA-2017 benchmark, outperforming many existing adversarial-learning-based methods.
Xin Luo 0009, Wei Chen 0009, Zhengfa Liang, Chen Li 0034, Yusong Tan
Neural Networks3
2022 Parallax Attention for Unsupervised Stereo Correspondence Learning
abstract
Stereo image pairs encode 3D scene cues into stereo correspondences between the left and right images. To exploit 3D cues within stereo images, recent CNN based methods commonly use cost volume techniques to capture stereo correspondence over large disparities. However, since disparities can vary significantly for stereo cameras with different baselines, focal lengths and resolutions, the fixed maximum disparity used in cost volume techniques hinders them to handle different stereo image pairs with large disparity variations. In this paper, we propose a generic parallax-attention mechanism (PAM) to capture stereo correspondence regardless of disparity variations. Our PAM integrates epipolar constraints with attention mechanism to calculate feature similarities along the epipolar line to capture stereo correspondence. Based on our PAM, we propose a parallax-attention stereo matching network (PASMnet) and a parallax-attention stereo image super-resolution network (PASSRnet) for stereo matching and stereo image super-resolution tasks. Moreover, we introduce a new and large-scale dataset named Flickr1024 for stereo image super-resolution. Experimental results show that our PAM is generic and can effectively learn stereo correspondence under large disparity variations in an unsupervised manner. Comparative results show that our PASMnet and PASSRnet achieve the state-of-the-art performance.
Longguang Wang, Yulan Guo, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Multi-Scale Cascade Disparity Refinement Stereo Network
abstract
Stereo matching has attracted much attention in recent years. Traditional methods can quickly generate a disparity result, but the accuracy is low. On the contrary, methods based on neural networks can achieve a high accuracy level, but they are difficult to reach the real-time level. Therefore, this paper presents MCDRNet, which combines traditional methods with neural networks to achieve real-time and accurate stereo matching results. Concretely, our network first generates a rough disparity map based on the traditional ADCensus algorithm. Then we design a novel Multi-Scale Cascade Network to refine the disparity map from coarse to fine. We evaluate our best-trained model on the KITTI official website. The results show that our network is much faster than most current top-performing methods(31×than CSPN, 56×than GANet, etc.). Meanwhile, it is more accurate than traditional stereo methods(SGM, SPS-St) and other fast 2D convolution networks(Fast DS-CS, DispNetC, etc.), demonstrating the rationalities and feasibilities of our method.
Xiaogang Jia, Wei Chen 0009, Zhengfa Liang, Xin Luo 0009, Mingfei Wu, Yusong Tan, Libo Huang 0002
ICASSP3
2021 Multi-Scale Cost Volumes Cascade Network for Stereo Matching
abstract
Stereo matching is essential for robot navigation. However, the accuracy of current widely used traditional methods is low, while methods based on CNN need expensive computational cost and running time. This is because different cost volumes play a crucial role in balancing speed and accuracy. Thus we propose MSCVNet, which combines traditional methods and neural networks to improve the quality of cost volume. Concretely, our network first generates multiple 3D cost volumes with different resolutions and then uses 2D convolutions to construct a novel cascade hourglass network for cost aggregation. Meanwhile, we design an algorithm to distinguish and calculate the loss for discontinuous areas of the disparity result. According to the KITTI official website, our network is much faster than most top-performing methods (24than CSPN, 44than GANet, etc.). Meanwhile, compared to traditional methods (SPS-St, SGM) and other real-time stereo matching networks (Fast DS-CS, DispNetC, and RTSNet, etc.), our network achieves a big improvement in accuracy, demonstrating the feasibility and capability of the proposed method.
Xiaogang Jia, Wei Chen 0009, Chen Li 0034, Zhengfa Liang, Mingfei Wu, Yusong Tan, Libo Huang 0002
ICRA4
2021 Fast and Accurate Lane Detection via Frequency Domain Learning
abstract
It is desirable to maintain both high accuracy and runtime efficiency in lane detection. State-of-the-art methods mainly address the efficiency problem by direct compression of high-dimensional features. These methods usually suffer from information loss and cannot achieve satisfactory accuracy performance. To ensure the diversity of features and subsequently maintain information as much as possible, we introduce multi-frequency analysis into lane detection. Specifically, we propose a multi-spectral feature compressor (MSFC) based on two-dimensional (2D) discrete cosine transform (DCT) to compress features while preserving diversity information. We group features and associate each group with an individual frequency component, which incurs only 1/7 overhead of one-dimensional convolution operation but preserves more information. Moreover, to further enhance the discriminability of features, we design a multi-spectral lane feature aggregator (MSFA) based on one-dimensional (1D) DCT to aggregate features from each lane according to their corresponding frequency components. The proposed method outperforms the state-of-the-art methods (including LaneATT and UFLD) on TuSimple, CULane, and LLAMAS benchmarks. For example, our method achieves 76.32% F1 at 237 FPS and 76.98% F1 at 164 FPS on CULane, which is 1.23% and 0.30% higher than LaneATT. Our code and models are available at https://github.com/harrylin-hyl/MSLD.
Yu-Lin He, Wei Chen 0009, Zhengfa Liang, Dan Chen 0001, Yusong Tan, Xin Luo 0009, Chen Li 0034, Yulan Guo
ACM Multimedia3
2021 Stereo Matching Using Multi-Level Cost Volume and Multi-Scale Feature Constancy
abstract
For CNNs based stereo matching methods, cost volumes play an important role in achieving good matching accuracy. In this paper, we present an end-to-end trainable convolution neural network to fully use cost volumes for stereo matching. Our network consists of three sub-modules, i.e., shared feature extraction, initial disparity estimation, and disparity refinement. Cost volumes are calculated at multiple levels using the shared features, and are used in both initial disparity estimation and disparity refinement sub-modules. To improve the efficiency of disparity refinement, multi-scale feature constancy is introduced to measure the correctness of the initial disparity in feature space. These sub-modules of our network are tightly-coupled, making it compact and easy to train. Moreover, we investigate the problem of developing a robust model to perform well across multiple datasets with different characteristics. We achieve this by introducing a two-stage finetuning scheme to gently transfer the model to target datasets. Specifically, in the first stage, the model is finetuned using both a large synthetic dataset and the target datasets with a relatively large learning rate, while in the second stage the model is trained using only the target datasets with a small learning rate. The proposed method is tested on several benchmarks including the Middlebury 2014, KITTI 2015, ETH3D 2017, and SceneFlow datasets. Experimental results show that our method achieves the state-of-the-art performance on all the datasets. The proposed method also won the 1st prize on the Stereo task of Robust Vision Challenge 2018.
Zhengfa Liang, Yulan Guo, Yiliu Feng, Wei Chen 0009, Linbo Qiao, Li Zhou 0009, Hengzhu Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Multi-Level Context Ultra-Aggregation for Stereo Matching
abstract
Exploiting multi-level context information to cost volume can improve the performance of learning-based stereo matching methods. In recent years, 3-D Convolution Neural Networks (3-D CNNs) show the advantages in regularizing cost volume but are limited by unary features learning in matching cost computation. However, existing methods only use features from plain convolution layers or a simple aggregation of multi-level features to calculate cost volume, which is insufficient because stereo matching requires discriminative features to identify corresponding pixels in rectified stereo image pairs. In this paper, we propose a unary features descriptor using multi-level context ultra-aggregation (MCUA), which encapsulates all convolutional features into a more discriminative representation by intra- and inter-level features combination. Specifically, a child module that takes low-resolution images as input captures larger context information; the larger context information from each layer is densely connected to the main branch of the network. MCUA makes good usage of multi-level features with richer context and performs the image-to-image prediction holistically. We introduce our MCUA scheme for cost volume calculation and test it on PSM-Net. We also evaluate our method on Scene Flow and KITTI 2012/2015 stereo datasets. Experimental results show that our method outperforms state-of-the-art methods by a notable margin and effectively improves the accuracy of stereo matching.
Guang-Yu Nie, Ming-Ming Cheng, Yun Liu 0011, Zhengfa Liang, Deng-Ping Fan, Yue Liu 0005, Yongtian Wang
CVPR4
2019 Learning Parallax Attention for Stereo Image Super-Resolution
abstract
Stereo image pairs can be used to improve the performance of super-resolution (SR) since additional information is provided from a second viewpoint. However, it is challenging to incorporate this information for SR since disparities between stereo images vary significantly. In this paper, we propose a parallax-attention stereo superresolution network (PASSRnet) to integrate the information from a stereo image pair for SR. Specifically, we introduce a parallax-attention mechanism with a global receptive field along the epipolar line to handle different stereo images with large disparity variations. We also propose a new and the largest dataset for stereo image SR (namely, Flickr1024). Extensive experiments demonstrate that the parallax-attention mechanism can capture correspondence between stereo images to improve SR performance with a small computational and memory cost. Comparative results show that our PASSRnet achieves the state-of-the-art performance on the Middlebury, KITTI 2012 and KITTI 2015 datasets.
Longguang Wang, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo
CVPR3
2018 Learning for Disparity Estimation Through Feature Constancy
abstract
Stereo matching algorithms usually consist of four steps, including matching cost calculation, matching cost aggregation, disparity calculation, and disparity refinement. Existing CNN-based methods only adopt CNN to solve parts of the four steps, or use different networks to deal with different steps, making them difficult to obtain the overall optimal solution. In this paper, we propose a network architecture to incorporate all steps of stereo matching. The network consists of three parts. The first part calculates the multi-scale shared features. The second part performs matching cost calculation, matching cost aggregation and disparity calculation to estimate the initial disparity using shared features. The initial disparity and the shared features are used to calculate the feature constancy that measures correctness of the correspondence between two input images. The initial disparity and the feature constancy are then fed into a sub-network to refine the initial disparity. The proposed method has been evaluated on the Scene Flow and KITTI datasets. It achieves the state-of-the-art performance on the KITTI 2012 and KITTI 2015 benchmarks while maintaining a very fast running time. Source code is available at http://github.com/leonzfa/iResNet.
Zhengfa Liang, Yiliu Feng, Yulan Guo, Hengzhu Liu, Wei Chen 0009, Linbo Qiao, Li Zhou 0009
CVPR1
2017 Learning Generic Features from DEH Channels for Object Proposals Generation
abstract
Depth information plays an important role in the human visual system, however it is not yet well- explored in existing proposal generation models. In this paper we propose a new geocentric embedding for depth images that encodes depth to the camera, the structural edges and height above ground-plane for each pixel named DEH channels. We demonstrate that this geocentric embedding works can be use to generate high quality object proposals with convolutional neural networks. Our experiments show significant performance improvements over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. We also exploit Fast R- CNN [8] on top of these proposals to perform object detection with RGB images, our approach obtains state-of-the-art results on all three KITTI object classes.
Yiliu Feng, Zhengfa Liang, Hengzhu Liu
SMARTCOMP2
2016 A refinement framework for background subtraction based on color and depth data
abstract
We present a refinement framework for background subtraction based on color and depth data. The foreground objects are segmented based on color and depth data independently, in which all of the existed background subtraction (BGS) methods can be applied. The two detected foregrounds will be very inaccurate in some situations such as shadowing and color camouflage. We focus our works on refining the inaccurate results by a supervised learning way. We propose to re-extract features from the source color and depth data. The features together with the initial detection results are fed to classifiers to obtain a better foreground detection. Experiments show that our method can take full advantage of the both information to detect foreground in color camouflage and shadowing situations, giving a promising result which is robust to the inaccurate initial detections, and outperforming the state-of-art algorithms that based on color and depth data.
Zhengfa Liang, Hengzhu Liu
ICIP1
2015 Simulation study of N-hit SET variation in differential cascade voltage switch logical circuits
Pengcheng Huang 0003, Shuming Chen, Zhengfa Liang, Chunmei Hu, Biwei Liu
Sci. China Inf. Sci.5