Zhong Liu 0005

dblp:30/2371-5 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0003-3242-5997ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Learnable patchmatch and self-teaching for multi-frame depth estimation in monocular endoscopy
Shuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu, Zhong Liu 0005
Eng. Appl. Artif. Intell.5
2025 Efficient motion feature aggregation for optical flow via locality-sensitive hashing
Weihai Chen, Xingming Wu, Zhong Liu 0005, Zhengguo Li
Neurocomputing4
2025 CrossFlow: Learning cost volumes for optical flow by cross-matching local and non-local image features
Zimeng Liu, Xingming Wu, Weihai Chen, Zhong Liu 0005, Zhengguo Li
J. Vis. Commun. Image Represent.5
2024 Online Unsupervised Video Object Segmentation via Contrastive Motion Clustering
abstract
Online unsupervised video object segmentation (UVOS) uses the previous frames as its input to automatically separate the primary object(s) from a streaming video without using any further manual annotation. A major challenge is that the model has no access to the future and must rely solely on the history, i.e., the segmentation mask is predicted from the current frame as soon as it is captured. In this work, a novel contrastive motion clustering algorithm with an optical flow as its input is proposed for the online UVOS by exploiting the common fate principle that visual elements tend to be perceived as a group if they possess the same motion pattern. We build a simple and effective auto-encoder to iteratively summarize non-learnable prototypical bases for the motion pattern, while the bases in turn help learn the representation of the embedding network. Further, a contrastive learning strategy based on a boundary prior is developed to improve foreground and background feature discrimination in the representation learning stage. The proposed algorithm can be optimized on arbitrarily-scale data (i.e., frame, clip, dataset) and performed in an online fashion. Experiments on$\textit {DAVIS}_{\textit {16}}$, FBMS, and SegTrackV2 datasets show that the accuracy of our method surpasses the previous state-of-the-art (SoTA) online UVOS method by a margin of 0.8%, 2.9%, and 1.1%, respectively. Furthermore, by using an online deep subspace clustering to tackle the motion grouping, our method is able to achieve higher accuracy at$3\times $faster inference time compared to SoTA online UVOS method, and making a good trade-off between effectiveness and efficiency. Our code is available athttps://github.com/xilin1991/CluterNet.
Lin Xi, Weihai Chen, Xingming Wu, Zhong Liu 0005, Zhengguo Li
IEEE Trans. Circuits Syst. Video Technol.4
2024 URCDC-Depth: Uncertainty Rectified Cross-Distillation With CutFlip for Monocular Depth Estimation
abstract
This work aims to estimate a high-quality depth map from a single RGB image. Due to the lack of depth clues, making full use of the long-range correlation and local information is critical for accurate depth estimation. To this end, we introduce an uncertainty rectified cross-distillation between the Transformer and convolutional neural network (CNN) to achieve a comprehensive depth estimator. Specifically, we utilize the depth estimates from the Transformer branch and CNN branch as pseudo labels to teach each other. At the same time, the pixel-wise depth uncertainty is modeled to mitigate the negative impact of noisy pseudo labels. To avoid the large capacity gap induced by the strong Transformer branch deteriorating the cross-distillation, we transfer the feature maps from the Transformer to the CNN and develop coupling units to assist the weak CNN branch in leveraging the transferred features. Furthermore, we introduce CutFlip, a surprisingly simple yet highly effective data augmentation technique, which forces the model to focus on more valuable depth reasoning clues apart from the vertical image position. Extensive experiments demonstrate that our model, termedURCDC-Depth, exceeds in performance previous state-of-the-art approaches on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets, with no additional computational burden in the evaluation phase. The source code will be publicly available upon acceptance. The source code is available athttps://github.com/ShuweiShao/URCDC-Depth.
Shuwei Shao, Zhongcai Pei, Weihai Chen, Zhong Liu 0005, Zhengguo Li
IEEE Trans. Multim.5
2023 Monocular Depth Estimation: A Survey
abstract
Monocular depth estimation is an ill-posed task in computer vision, which holds great significance in the fields such as artificial intelligence, virtual reality, augmented reality, path planning, unmanned driving, and navigation guidance. The primary objective of monocular depth estimation is to predict the depth value of each pixel or infer depth information, given just a single red-green-blue (RGB) image as input. Traditional monocular depth estimation methods rely on limited depth cues, such as strict scene conditions. With the significant advancements in computer vision and artificial intelligence, monocular depth estimation using deep learning has been extensively researched and has yielded substantial results. This paper presents a comprehensive survey of monocular depth estimation. Firstly, we give an overall introduction to monocular depth estimation and explain it from traditional and deep learning-based methods, respectively. To specify, supervised, self-supervised and semi-supervised models are described in detail in deep learning-based methods. Additionally, we introduce publicly available benchmark datasets and evaluation metrics commonly used in this field. Finally, we discuss the current challenges and promising prospects for the development of monocular depth estimation.
Dong Wang 0051, Zhong Liu 0005, Shuwei Shao, Xingming Wu, Weihai Chen, Zhengguo Li
IECON2
2023 IEBins: Iterative Elastic Bins for Monocular Depth Estimation
abstract
Monocular depth estimation (MDE) is a fundamental topic of geometric computer vision and a core technique for many downstream applications. Recently, several methods reframe the MDE as a classification-regression problem where a linear combination of probabilistic distribution and bin centers is used to predict depth. In this paper, we propose a novel concept of iterative elastic bins (IEBins) for the classification-regression-based MDE. The proposed IEBins aims to search for high-quality depth by progressively optimizing the search range, which involves multiple stages and each stage performs a finer-grained depth search in the target bin on top of its previous stage. To alleviate the possible error accumulation during the iterative process, we utilize a novel elastic target bin to replace the original target bin, the width of which is adjusted elastically based on the depth uncertainty. Furthermore, we develop a dedicated framework composed of a feature extractor and an iterative optimizer that has powerful temporal context modeling capabilities benefiting from the GRU-based architecture. Extensive experiments on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets demonstrate that the proposed method surpasses prior state-of-the-art competitors. The source code is publicly available at https://github.com/ShuweiShao/IEBins.
Shuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu 0005, Weihai Chen, Zhengguo Li
NeurIPS4
2023 Unsupervised Optical Flow Estimation for Differently Exposed Images in LDR Domain
abstract
Differently exposed low dynamic range (LDR) images are often captured sequentially using a smart phone or a digital camera with movements. Optical flow thus plays an important role in ghost removal for high dynamic range (HDR) imaging. The optical flow estimation is based on the theory of photometric consistency, which assumes that the corresponding pixels between two images have the same intensity. However, the assumption is no longer valid for the differently exposed LDR images since a pixel’s intensity changes significantly inter images. To address the problem, an unsupervised optical flow estimation framework, is presented in this study. Intensity mapping functions (IMFs) are first adopted to alleviate the intensity changes between the LDR images. Then a novel IMF-based unsupervised learning objective is proposed to circumvent the need for ground truth optical flows when training the deep network. Experimental results and ablation studies on publicly available datasets show that our framework outperforms the state-of-the-art unsupervised optical flow methods, demonstrating the effectiveness of the IMF and the learning objective. Our code is available athttps://github.com/liuziyang123/LDRFlow.
Zhengguo Li, Weihai Chen, Xingming Wu, Zhong Liu 0005
IEEE Trans. Circuits Syst. Video Technol.5
2023 Self-Supervised Monocular Depth Estimation With Self-Reference Distillation and Disparity Offset Refinement
abstract
Monocular depth estimation plays a fundamental role in computer vision. Due to the costly acquisition of depth ground truth, self-supervised methods that leverage adjacent frames to establish a supervision signal have emerged as the most promising paradigms. In this work, we propose two novel ideas to improve self-supervised monocular depth estimation: 1) self-reference distillation and 2) disparity offset refinement. Specifically, we use a parameter-optimized model as the teacher updated as the training epochs to provide additional supervision during the training process. The teacher model has the same structure as the student model, with weights inherited from the historical student model. In addition, a multiview check is introduced to filter out the outliers produced by the teacher model. Furthermore, we leverage the contextual consistency between high-level and low-level features to obtain multiscale disparity offsets, which are used to refine the disparity output incrementally by aligning disparity information at different scales. The experimental results on the KITTI and Make3D datasets show that our method outperforms previous state-of-the-art competitors.
Zhong Liu 0005, Shuwei Shao, Xingming Wu, Weihai Chen
IEEE Trans. Circuits Syst. Video Technol.1
2023 Towards Comprehensive Monocular Depth Estimation: Multiple Heads are Better Than One
abstract
Depth estimation attracts widespread attention in the computer vision community. However, it is still quite difficult to recover an accurate depth map using only one RGB image. We observe a phenomenon that existing methods tend to fail in different cases, caused by differences in network architecture, loss function and so on. In this work, we investigate into the phenomenon and propose to integrate the strengths of multiple weak depth predictor to build a comprehensive and accurate depth predictor, which is critical for many real-world applications, e.g., 3D reconstruction. Specifically, we construct multiple base (weak) depth predictors by utilizing different Transformer-based and convolutional neural network (CNN)-based architectures. Transformer establishes long-range correlation while CNN preserves local information ignored by Transformer due to the spatial inductive bias. Therefore, the coupling of Transformer and CNN contributes to the generation of complementary depth estimates, which are essential to achieve a comprehensive depth predictor. Then, we design mixers to learn from multiple weak predictions and adaptively fuse them into a strong depth estimate. The resultant model, which we refer to as Transformer-assisted depth ensembles (TEDepth). On the standard NYU-Depth-v2 and KITTI datasets, we thoroughly explore how the neural ensembles affect the depth estimation and demonstrate that our TEDepth achieves better results than previous state-of-the-art approaches. To validate the generalizability across cameras, we directly apply the models trained on NYU-Depth-v2 to the SUN RGB-D dataset without any fine-tuning, and the superior results emphasize its strong generalizability.
Shuwei Shao, Zhongcai Pei, Zhong Liu 0005, Weihai Chen, Wentao Zhu 0001, Xingming Wu, Baochang Zhang 0001
IEEE Trans. Multim.4
2022 SEHLNet: Separate Estimation of High- and Low-Frequency components for Depth Completion
abstract
Depth completion refers to inferring the dense depth map from a sparse depth map with or without corre-sponding color image. Numerous neural networks have been proposed to accomplish this task. However, insufficient uti-lization of heteromorphic data and the fact that predicted dense depth prefers a sparse depth enormously damage the performance of approaches. To reduce data preference and fully utilize two modalities, this paper proposes a novel network that predicts high- and low-frequency components of dense depth separately. Specifically, the framework consists of a Low-Frequency(LF) branch and a High-Frequency(HF) branch. In the LF branch, we recover the low-frequency depth component from sparse depth through an Adaptive Graph-Generate Graph Attention Network, which can be seen as a low-pass filter. In the HF branch, we model the high-frequency component, e.g. boundaries, as residuals to mitigate the impact of data preferences. Moreover, in this branch, we propose an Attention-based Self-Fusion mechanism to efficiently fuse multi-scale features extracted from the sparse depth and color image. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on the KITTI benchmark and ranks 1st in root mean squared error among other published approaches.
Haosong Yue, Zhanggang Lyu, Wei Wang 0036, Zhong Liu 0005, Weihai Chen
ICRA5
2022 Discriminative and semantic feature selection for place recognition towards dynamic environments
Jinyu Miao, Xingming Wu, Haosong Yue, Zhong Liu 0005, Weihai Chen
Pattern Recognit. Lett.5
2022 DSRGAN: Detail Prior-Assisted Perceptual Single Image Super-Resolution via Generative Adversarial Networks
abstract
The generative adversarial network (GAN) is successfully applied to study the perceptual single image super-resolution (SISR). However, since the GAN is data-driven, it has a fundamental limitation on restoring real high frequency information for an unknown instance (or image) during test. On the other hand, the conventional model-based methods have a superiority to achieve instance adaptation as they operate by considering the statistics of each instance (or image) only. Motivated by this, we propose a novel model-based algorithm, which can extract the detail layer of an image efficiently. The detail layer represents the high frequency information of image and it is constituted of image edges and fine textures. It is seamlessly incorporated into the GAN and serves as a prior knowledge to assist the GAN in generating more realistic details. The proposed method, named DSRGAN, takes advantages from both the model-based conventional algorithm and the data-driven deep learning network. Experimental results demonstrate that the DSRGAN outperforms the state-of-the-art SISR methods on perceptual metrics, meanwhile achieving comparable results in terms of fidelity metrics. Following the DSRGAN, it is feasible to incorporate other conventional image processing algorithms into a deep learning network to form a model-based deep SISR.
Zhengguo Li, Xingming Wu, Zhong Liu 0005, Weihai Chen
IEEE Trans. Circuits Syst. Video Technol.4
2022 Implicit Motion-Compensated Network for Unsupervised Video Object Segmentation
abstract
Unsupervised video object segmentation (UVOS) aims at automatically separating the primary foreground object(s) from the background in a video sequence. Existing UVOS methods either lack robustness when there are visually similar surroundings (appearance-based) or suffer from deterioration in the quality of their predictions because of dynamic background and inaccurate flow (flow-based). To overcome the limitations, we propose an implicit motion-compensated network (IMCNet) combining complementary cues (i.e., appearance and motion) with aligned motion information from the adjacent frames to the current frame at the feature level without estimating optical flows. The proposed IMCNet consists of an affinity computing module (ACM), an attention propagation module (APM), and a motion compensation module (MCM). The light-weight ACM extracts commonality between neighboring input frames based on appearance features. The APM then transmits global correlation in a top-down manner. Through coarse-to-fine iterative inspiring, the APM will refine object regions from multiple resolutions so as to efficiently avoid losing details. Finally, the MCM aligns motion information from temporally adjacent frames to the current frame which achieves implicit motion compensation at the feature level. We perform extensive experiments on$\textit {DAVIS}_{\textit {16}}$and$\textit {YouTube-Objects}$. Our network achieves favorable performance while running at a faster speed compared to the state-of-the-art methods. Our code is available athttps://github.com/xilin1991/IMCNet.
Lin Xi, Weihai Chen, Xingming Wu, Zhong Liu 0005, Zhengguo Li
IEEE Trans. Circuits Syst. Video Technol.4
2014 Salient region detection using high level feature
abstract
In the last few decades, selective visual attention has been extensively studied for its promising contributions to computer vision applications. Many different models have been proposed to compute visual saliency, which can be coarsely formulated as computational or psychophysical. Most existing methods are based on bottom-up mechanism, an automatic human behavior to guide gaze allocation. And low level features such as color, intensity and orientation are commonly adopted to compute saliency map. In this work, we propose a saliency computation method that integrates high-level information of object with low-level features. The result map is more suitable for most top-down tasks in the field of mobile robot requiring object information.
Zhong Liu 0005, Weihai Chen, Xingming Wu
ICARCV1
2012 Novel Spatial Pyramid Matching for scene and object classification
abstract
It is difficult to classify object or scene images with high accuracy when the dataset is relatively large. Spatial Pyramid Matching (SPM) was proposed to deal with this problem, but there are some shortages. As an improvement for SPM, we proposed three pieces of meliorations: first, use approximate nearest neighbor method instead of k-means for clustering; second, regulate the size of codebook referring to quantity and pixels of the images, by calculating sub-codebook for every category and eliminating the codes which are nearer to the registered ones than the threshold; third, rescale the histogram features, and classify the scene with hierarchical strategy. Experiments prove that our approach make better performance than other state-of-the-art classification methods using just one matching kernel.
Weihai Chen, Xingming Wu, Zhong Liu 0005
INDIN4
2012 Regions of interest extraction based on HSV color space
abstract
In this paper, a simple method to extract regions of interest (ROI) from images is proposed. In the field of image processing, intensity, color and orientation are commonly used features for saliency map generation in most visual attention model. However, texture feature can contribute to the guidance of attention in a bottom-up model. We consider texture contrast as a component of final saliency map. Hue, saturation, and value (HSV) color space is also adopted in this paper for its good capability of representing the colors of human perception and simplicity of computation. Moreover, binocular stereo image pair is adopted as source image. The result shows that the proposed saliency computational method can effectively detect salient region, and it is more suitable for environmental perception and cognition, object detection, and mobile robot navigation.
Zhong Liu 0005, Weihai Chen, Yuhua Zou, Cun Hu
INDIN1
2012 Indoor localization and 3D scene reconstruction for mobile robots using the Microsoft Kinect sensor
abstract
In this paper we present an approach to indoor localization and 3D scene reconstruction using the Microsoft Kinect sensor. The proposed system can simultaneously estimates the position and orientation of a hand-held Kinect and generates a dense 3D model of the indoor environment. Furthermore, the robustness and processing time for four different feature descriptors (SURF, ORB, Shi-Tomasi and FAST) are evaluated. The experiment results demonstrate that our system can robustly deal with complicated data in common indoor scenarios while running in semi-real-time.
Yuhua Zou, Weihai Chen, Xingming Wu, Zhong Liu 0005
INDIN4