VLDB 2026 Research / reviewers in the wild / expert
Qi Liu 0054
dblp:95/2446-54
· DBLP profile ↗
13ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-4974-1518ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-event representation and multi-level fusion for robust RGB-event object tracking
Bin Fan 0002, Zhexiong Wan, Qi Liu 0054, Yuchao Dai |
Knowl. Based Syst. | 4 |
| 2025 | Geometry-Aware 3D Salient Object Detection NetworkabstractPoint cloud salient object detection has attracted the attention of researchers in recent years. Since existing works do not fully utilize the geometry context of 3D objects, blurry boundaries are generated when segmenting objects with complex backgrounds. In this paper, we propose a geometry-aware 3D salient object detection network that explicitly clusters points into superpoints to enhance the geometric boundaries of objects, thereby segmenting complete objects with clear boundaries. Specifically, we first propose a simple yet effective superpoint partition module to cluster points into superpoints. In order to improve the quality of superpoints, we present a point cloud class-agnostic loss to learn discriminative point features for clustering superpoints from the object. After obtaining superpoints, we then propose a geometry enhancement module that utilizes superpoint-point attention to aggregate geometric information into point features for predicting the salient map of the object with clear boundaries. Extensive experiments show that our method achieves new state-of-the-art performance on the PCSOD dataset. Chen Wang 0049, Le Hui, Qi Liu 0054, Yuchao Dai |
AAAI | 4 |
| 2024 | Improving Depth Completion via Depth Feature UpsamplingabstractThe encoder-decoder network (ED-Net) is a commonly employed choice for existing depth completion methods, but its working mechanism is ambiguous. In this paper, we vi-sualize the internal feature maps to analyze how the net-work densifies the input sparse depth. We find that the en-coder feature of ED-Net focus on the areas with input depth points around. To obtain a dense feature and thus esti-mate complete depth, the decoder feature tends to comple-ment and enhance the encoder feature by skip-connection to make the fused encoder-decoder feature dense, resulting in the decoder feature also exhibits sparse. However, ED-Net obtains the sparse decoder feature from the dense fused feature at the previous stage, where the “dense-i-sparse‘’ process destroys the completeness of features and loses in-formation. To address this issue, we present a depth feature upsampling network (DFU) that explicitly utilizes these dense features to guide the upsampling of a low-resolution (LR) depth feature to a high-resolution (HR) one. The completeness of features is maintained throughout the up-sampling process, thus avoiding information loss. Fur-thermore, we propose a confidence-aware guidance module (CGM), which is confidence-aware and performs guidance with adaptive receptive fields (GARF), to fully exploit the potential of these dense features as guidance. Experimental results show that our DFU, a plug-and-play module, can significantly improve the performance of existing ED-Net based methods with limited computational overheads, and new SOTA results are achieved. Besides, the generalization capability on sparser depth is also enhanced. Project page: https://npucvr.github.iolDFU. Ge Zhang 0006, Shaoqian Wang, Bo Li 0090, Qi Liu 0054, Le Hui, Yuchao Dai |
CVPR | 5 |
| 2024 | 3D Focusing-and-Matching Network for Multi-Instance Point Cloud RegistrationabstractMulti-instance point cloud registration aims to estimate the pose of all instances of a model point cloud in the whole scene. Existing methods all adopt the strategy of first obtaining the global correspondence and then clustering to obtain the pose of each instance. However, due to the cluttered and occluded objects in the scene, it is difficult to obtain an accurate correspondence between the model point cloud and all instances in the scene. To this end, we propose a simple yet powerful 3D focusing-and-matching network for multi-instance point cloud registration by learning the multiple pair-wise point cloud registration. Specifically, we first present a 3D multi-object focusing module to locate the center of each object and generate object proposals. By using self-attention and cross-attention to associate the model point cloud with structurally similar objects, we can locate potential matching instances by regressing object centers. Then, we propose a 3D dual-masking instance matching module to estimate the pose between the model point cloud and each object proposal. It performs instance mask and overlap mask masks to accurately predict the pair-wise correspondence. Extensive experiments on two public benchmarks, Scan2CAD and ROBI, show that our method achieves a new state-of-the-art performance on the multi-instance point cloud registration task. Le Hui, Qi Liu 0054, Bo Li 0090, Yuchao Dai |
NeurIPS | 3 |
| 2024 | Decomposed Guided Dynamic Filters for Efficient RGB-Guided Depth CompletionabstractRGB-guided depth completion aims at predicting dense depth maps from sparse depth measurements and corresponding RGB images, where how to effectively and efficiently exploit the multi-modal information is a key issue. Guided dynamic filters, which generate spatially-variant depth-wise separable convolutional filters from RGB features to guide depth features, have been proven to be effective in this task. However, the dynamically generated filters require massive model parameters, computational costs and memory footprints when the number of feature channels is large. In this paper, we propose to decompose the guided dynamic filters into a spatially-shared component multiplied by content-adaptive adaptors at each spatial location. Based on the proposed idea, we introduce two decomposition schemes$\mathcal {A}$and$\mathcal {B}$, which decompose the filters by splitting the filter structure and using spatial-wise attention, respectively. The decomposed filters not only maintain the favorable properties of guided dynamic filters as being content-dependent and spatially-variant, but also reduce model parameters and hardware costs, as the learned adaptors are decoupled with the number of feature channels. Extensive experimental results demonstrate that the methods using our schemes outperform state-of-the-art methods on the KITTI dataset, and rank 1st and 2nd on the KITTI benchmark at the time of submission. Meanwhile, they also achieve comparable performance on the NYUv2 dataset. In addition, our proposed methods are general and could be employed as plug-and-play feature fusion blocks in other multi-modal fusion tasks such as RGB-D salient object detection. Yuxin Mao, Qi Liu 0054, Yuchao Dai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Joint Appearance and Motion Learning for Efficient Rolling Shutter CorrectionabstractRolling shutter correction (RSC) is becoming increasingly popular for RS cameras that are widely used in commercial and industrial applications. Despite the promising performance, existing RSC methods typically employ a two-stage network structure that ignores intrinsic infor-mation interactions and hinders fast inference. In this pa-per, we propose a single-stage encoder-decoder-based network, named JAMNet, for efficient RSC. It first extracts pyramid features from consecutive RS inputs, and then simultaneously refines the two complementary information (i.e., global shutter appearance and undistortion motion field) to achieve mutual promotion in a joint learning de-coder. To inject sufficient motion cues for guiding joint learning, we introduce a transformer-based motion embed-ding module and propose to pass hidden states across pyra-mid levels. Moreover, we present a new data augmentation strategy “vertical flip + inverse order” to release the potential of the RSC datasets. Experiments on various benchmarks show that our approach surpasses the state-of-the-art methods by a large margin, especially with a 4.7 dB PSNR leap on real-world RSC. Code is available at https://github.com/GitCVfb/JAMNet. Bin Fan 0002, Yuxin Mao, Yuchao Dai, Zhexiong Wan, Qi Liu 0054 |
CVPR | 5 |
| 2023 | LRRU: Long-short Range Recurrent Updating Networks for Depth CompletionabstractExisting deep learning-based depth completion methods generally employ massive stacked layers to predict the dense depth map from sparse input data. Although such approaches greatly advance this task, their accompanied huge computational complexity hinders their practical applications. To accomplish depth completion more efficiently, we propose a novel lightweight deep network framework, the Long-short Range Recurrent Updating (LRRU) network. Without learning complex feature representations, LRRU first roughly fills the sparse input to obtain an initial dense depth map, and then iteratively updates it through learned spatially-variant kernels. Our iterative update process is content-adaptive and highly flexible, where the kernel weights are learned by jointly considering the guidance RGB images and the depth map to be updated, and large-to-small kernel scopes are dynamically adjusted to capture long-to-short range dependencies. Our initial depth map has coarse but complete scene depth information, which helps relieve the burden of directly regressing the dense depth from sparse ones, while our proposed method can effectively refine it to an accurate depth map with less learnable parameters and inference time. Experimental results demonstrate that our proposed LRRU variants achieve state-of-the-art performance across different parameter regimes. In particular, the LRRU-Base model outperforms competing approaches on the NYUv2 dataset, and ranks 1st on the KITTI depth completion benchmark at the time of submission. Project page: https://npucvr.github.io/LRRU/. Bo Li 0090, Ge Zhang 0006, Qi Liu 0054, Tao Gao 0001, Yuchao Dai |
ICCV | 4 |
| 2022 | Context-Aware Video Reconstruction for Rolling Shutter CamerasabstractWith the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also places a higher demand on realism. Existing solutions, using deep neural networks or optimization, achieve promising performance. However, these methods generate intermediate GS frames through image warping based on the RS model, which inevitably result in black holes and noticeable motion artifacts. In this paper, we alleviate these issues by proposing a context-aware GS video reconstruction architecture. It facilitates the advantages such as occlusion reasoning, motion compensation, and temporal abstraction. Specifically, we first estimate the bilateral motion field so that the pixels of the two RS frames are warped to a common GS frame accordingly. Then, a refinement scheme is proposed to guide the GS frame synthesis along with bilateral occlusion masks to produce high-fidelity GS video frames at arbitrary times. Furthermore, we derive an approximated bilateral motion field model, which can serve as an alternative to provide a simple but effective GS frame initialization for related tasks. Experiments on synthetic and real data show that our approach achieves superior performance over state-of-the-art methods in terms of objective metrics and subjective visual quality. Code is available at https://github.com/GitCVfb/CVR. Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Qi Liu 0054, Mingyi He |
CVPR | 4 |
| 2022 | Searching Dense Point Correspondences via Permutation Matrix LearningabstractAlthough 3D point cloud data has received widespread attentions as a general form of 3D signal expression, applying point clouds to the task of dense correspondence estimation between 3D shapes has not been investigated widely. Furthermore, even in the few existing 3D point cloud-based methods, an important and widely acknowledged principle,i.e. one-to-one matching, is usually ignored. In response, this paper presents a novel end-to-end learning-based method to estimate the dense correspondence of 3D point clouds, in which the problem of point matching is formulated as a zero-one assignment problem to achieve a permutation matching matrix to implement the one-to-one principle fundamentally. Note that the classical solutions of this assignment problem are always non-differentiable, which is fatal for deep learning frameworks. Thus we design a special matching module, which solves a doubly stochastic matrix at first and then projects this obtained approximate solution to the desired permutation matrix. Moreover, to guarantee end-to-end learning and the accuracy of the calculated loss, we calculate the loss from the learned permutation matrix but propagate the gradient to the doubly stochastic matrix directly which bypasses the permutation matrix during the backward propagation. Our method can be applied to both non-rigid and rigid 3D point cloud data and extensive experiments show that our method achieves state-of-the-art performance for dense correspondence learning.The code will be released. Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Qi Liu 0054 |
IEEE Signal Process. Lett. | 5 |
| 2018 | Dominant vanishing point detection in the wild with application in composition analysis
Xiaodan Zhang 0005, Xinbo Gao 0001, Wen Lu 0004, Lihuo He, Qi Liu 0054 |
Neurocomputing | 5 |
| 2018 | Single Image Dehazing With Depth-Aware Non-Local Total Variation RegularizationabstractSingle image dehazing can benefit many computer vision applications hence has attracted much more attention in recent years. However, it still remains a challenging task due to its double uncertainty of scene transmission and scene radiance. The existing image dehazing methods usually impair edges in the estimated transmission which leads to halo effects in the dehazing results. Besides, most existing methods suffer from noise and artifacts amplification in dense haze region after dehazing. To address these challenges, we propose a transmission adaptive regularized image recovery method for high quality single image dehazing. An initial transmission map is first obtained by a boundary constraint on the haze model. Then it is refined by applying a non-local total variation (NLTV) regularization to keep depth structures while smoothing excessive details. Noticing that the artifacts amplification effect depends on scene transmission, a transmission adaptive regularized recovery method based on NLTV is proposed to simultaneously suppress visual artifacts and preserve image details in the final dehazing result. An efficient alternating optimization algorithm is also proposed to solve the regularization model. Thorough experimental results demonstrate that the proposed method can effectively suppress visual artifacts for degraded hazy images, and yields high-quality results comparative to the state-of-the-art dehazing methods both quantitatively and qualitatively. Qi Liu 0054, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
IEEE Trans. Image Process. | 1 |
| 2017 | Haze removal for a single visible remote sensing image
Qi Liu 0054, Xinbo Gao 0001, Lihuo He, Wen Lu 0004 |
Signal Process. | 1 |
| 2016 | Fast image quality assessment via supervised iterative quantization method
Lihuo He, Di Wang 0011, Qi Liu 0054, Wen Lu 0004 |
Neurocomputing | 3 |