EDBT 2026 Demo / reviewers in the wild / expert
Wei Liu 0044
dblp:49/3283-44
· DBLP profile ↗
51ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0001-6351-9019ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 23 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic ScenesabstractSelf-supervised monocular depth estimation methods severely compromise accuracy in dynamic objects due to their static scene assumption. Existing approaches for dynamic scenes suffer from two critical shortcomings: 1) reliance on supervised segmentation models (requiring costly annotations) or computationally intensive multi-branch models to isolate moving objects, and 2) simple integration of 2D/3D motion flow without reliable supervision for dynamic objects. We propose AdaDepth, a two‑stage framework that jointly performs unsupervised scene decomposition and dynamic-aware depth learning. In the initial structural stage, our geometry-motion joint scene decomposition (GMoDecomp) module ensures the robust generation of a depth prior and simultaneously partitions the scene into multiple regions through the fusion of geometric and motion cues. In the region-adaptive refinement stage, we exploit the depth prior and decomposed regions to introduce motion-aware and geometry-consistent constraints, effectively improving depth estimation in dynamic scenes. AdaDepth achieves accurate depth prediction in highly dynamic scenes without relying on external labels or specialized segmentation models. Extensive experiments on KITTI, Cityscapes, and Waymo Open demonstrate its superiority over state-of-the-art approaches. Xuanang Gao, Xiongbin Wu, Zhiwei Ning, Zhonglong Zheng, Jie Yang 0002, Wei Liu 0044 |
AAAI | 7 |
| 2026 | Exploiting All Mamba Fusion for Efficient RGB-D TrackingabstractDespite the progress made through deep learning, existing Visual Object Tracking (VOT) frameworks struggle with real-world challenges. Recent approaches incorporate additional modalities like Depth, Thermal Infrared, and Language to enhance the robustness of VOT, particularly with the improvement of the depth sensor precision, facilitating RGB-D tracking. However, current RGB-D trackers often copy RGB tracking paradigms, leading to inefficiency due to two-stream architectures that fail to exploit heterogeneous features, and reliance on simplistic or large-parameter fusion methods. To address these challenges, we propose AMTrack, a one-stream RGB-D tracker leveraging Mamba's linear complexity for simultaneous feature extraction and two-stage cross-modal feature fusion. Our innovation also includes a low-parameter Multimodal Mix Mamba (3M) module, which optimizes deep feature fusion and reduces computational overhead. The advantage of the 3M module stems from our Multimodal State Space Model (MSSM), a multimodal feature interaction component reconstructed based on SSM. Experiments across multiple RGB-D tracking datasets indicate that AMTrack achieves superior performance with lower parameters and memory demands compared to state-of-the-arts. Ge Ying, Dawei Zhang 0002, Chengzhuan Yang, Wei Liu 0044, Sang-Woon Jeon, Hua Wang 0002, Changqin Huang, Zhonglong Zheng |
AAAI | 4 |
| 2026 | Inter-modality feature prediction through multimodal fusion for 3D shape defect detection
Mujtaba Asad, Waqar Azeem, Hafiz Tayyab Mustafa, Yuming Fang 0001, Jie Yang 0002, Yifan Zuo 0001, Wei Liu 0044 |
Neural Networks | 7 |
| 2026 | D2S-RSG-SSD: Dual Double-Sampling With Random Sub-Samples Generation for Self-Supervised Real Image DenoisingabstractRecent advances in self-supervised image denoising have highlighted the potential of Blind-Spot Networks (BSNs). However, existing methods suffer from three major limitations: (1) Their effectiveness in real-world scenarios is limited by strong assumptions, such as noise independence, which rarely hold in practice. (2) While sampling-based strategies can partially improve performance, BSNs inherently suffer from information loss caused by centroid masking, and removing the blind spot leads to noise overfitting, both of which hinder denoising performance. (3) Sampling-based methods often introduce checkerboard artifacts, yet existing studies typically overlook the fundamental differences between these artifacts and real noise. To address these issues, we propose a novel self-supervised denoising framework, Dual Double-Sampling with Random Sub-samples Generation (D2S-RSG-SSD). To address Limitation 1, we introduce a sampling-based framework that breaks noise dependence by combining Random Sub-samples Generation (RSG) with a cross-paired loss $\mathcal {L}_{RSG}$LRSG. RSG generates diverse sub-samples with inherent variance, referred to as sampling differences, which serve as natural perturbations to augment training data and disrupt spatial noise correlations. The proposed loss function ensures full utilization of these sub-samples while stabilizing optimization. To address Limitation 2, we propose a Dual Double-Sampling (D2S) strategy with fixed sampling patterns and a dual-branch architecture. This design reduces reliance on pixel-level information and leverages complementary features to mitigate both noise overfitting and information loss. A key advantage is its compatibility with various advanced denoising networks, lifting the constraint of using BSNs in self-supervised settings. Additionally, we introduce a fixed sub-image sampling strategy to prevent pattern collapse during inference and ensure stability. To address Limitation 3, we explicitly differentiate checkerboard artifacts from real noise and develop a dedicated artifact remover to correct pixel discontinuities caused by sampling-based operations. This design preserves fine image details while reducing over-smoothing. Experiments on benchmark real-noise datasets and self-captured noisy images demonstrate the robustness and generalizability of our framework, achieving better performance over existing methods. Xiao Liu 0022, Xiuya Shi, Yizhong Pan, Shuhang Gu, Wei Liu 0044, Chao Ren 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection With IoU Joint PredictionabstractMulti-modal methods based on camera and LiDAR sensors have garnered significant attention in the field of 3D detection. However, many prevalent works focus on single or partial stage fusion, leading to insufficient feature extraction and suboptimal performance. In this paper, we introduce a multi-stage cross-modal fusion 3D detection framework, termed CMF-IOU, to effectively address the challenge of aligning 3D spatial and 2D semantic information. Specifically, we first project the pixel information into 3D space via a depth completion network to get the pseudo points, which unifies the representation of the LiDAR and camera information. Then, a bilateral cross-view enhancement 3D backbone is designed to encode LiDAR points and pseudo points. The first sparse-to-distant (S2D) branch utilizes an encoder-decoder structure to reinforce the representation of sparse LiDAR points. The second residual view consistency (ResVC) branch is proposed to mitigate the influence of inaccurate pseudo points via both the 3D and 2D convolution processes. Subsequently, we introduce an iterative voxel-point aware fine grained pooling module, which captures the spatial information from LiDAR points and textural information from pseudo points in the proposal refinement stage. To achieve more precise refinement during iteration, an intersection over union (IoU) joint prediction branch integrated with a novel proposals generation technique is designed to preserve the bounding boxes with both high IoU and classification scores. Extensive experiments show the superior performance of our method on the KITTI, nuScenes and Waymo datasets. The code is available at https://github.com/pami-zwning/CMF-IOU. Zhiwei Ning, Zhaojiang Liu, Xuanang Gao, Yifan Zuo 0001, Jie Yang 0002, Yuming Fang 0001, Wei Liu 0044 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | EDVD: Cross-Modal Spatio-Temporal Fusion With Event and Diffusion for Video DeblurringabstractRestoring high-quality images from blurred videos is a highly challenging task, especially in severely blurred scenes. In recent years, event-based methods have achieved significant progress in video deblurring. However, the modal differences between the event and image increase the difficulty of feature fusion. Additionally, the sparsity of event makes it difficult to restore some local details. To address these issues, we propose a new video deblurring method. Firstly, we design a cross-modal collaborative attention mechanism to effectively fuse features from blurred frames and event frames, thereby deeply extracting motion information from event frames. Secondly, we utilize a diffusion model to generate spatial guiding prior feature, enhancing local details and textures. Furthermore, we propose an event-guided dynamic feature fusion module that adaptively integrates spatio-temporal information from neighboring frames. Experimental results on both synthetic and real datasets demonstrate that our method outperforms the current state-of-the-art approaches. The code is available at: https://github.com/Frank-Zhou-01/EDVD-main. Ying Fu 0003, Tao Wu 0010, Qing Li 0001, Xi Wu 0004, Wei Liu 0044 |
IEEE Trans. Image Process. | 7 |
| 2025 | Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentationabstractIn clinical practice, medical image analysis often requires efficient execution on resource-constrained mobile devices. However, existing mobile models-primarily optimized for natural images-tend to perform poorly on medical tasks due to the significant information density gap between natural and medical domains. Combining computational efficiency with medical imaging-specific architectural advantages remains a challenge when developing lightweight, universal, and high-performing networks. To address this, we propose a mobile model called Mobile U-shaped Vision Transformer (Mobile U-ViT) tailored for medical image segmentation. Specifically, we employ the newly proposed ConvUtr as a hierarchical patch embedding, featuring a parameter-efficient large-kernel CNN with inverted bottleneck fusion. This design exhibits transformer-like representation learning capacity while being lighter and faster. To enable efficient local-global information exchange, we introduce a novel Large-kernel Local-Global-Local (LKLGL) block that effectively balances the low information density and high-level semantic discrepancy of medical images. Finally, we incorporate a shallow and lightweight transformer bottleneck for long-range modeling and employ a cascaded decoder with downsampled skip connections for dense prediction. Despite its reduced computational demands, our medical-optimized architecture achieves state-of-the-art performance across eight public 2D and 3D datasets covering diverse imaging modalities, including zero-shot testing on four unseen datasets. These results establish it as an efficient yet powerful and generalization solution for mobile medical image analysis. Code is available at: https://github.com/FengheTan9/Mobile-U-ViT. Fenghe Tang, Bingkun Nian, Jianrui Ding, Quan Quan, Chengqi Dong, Jie Yang 0002, Wei Liu 0044, Shaohua Kevin Zhou |
ACM Multimedia | 8 |
| 2025 | STDepth: Leveraging semantic-textural information in transformers for self-supervised monocular depth estimation
Xuanang Gao, Bingchao Wang, Zhiwei Ning, Jie Yang 0002, Wei Liu 0044 |
Comput. Vis. Image Underst. | 5 |
| 2025 | COTA-motion: Controllable image-to-video synthesis with dense semantic trajectories
Yirui Chen, Wenqing Chu, Jie Yang 0002, Xiaonan Mao, Wei Liu 0044 |
Neurocomputing | 6 |
| 2025 | MambaMIM: Pre-training Mamba with state space token interpolation and its application to medical image segmentation
Fenghe Tang, Bingkun Nian, Yingtai Li, Zihang Jiang, Jie Yang 0002, Wei Liu 0044, Shaohua Kevin Zhou |
Medical Image Anal. | 6 |
| 2025 | NAS-GS: Normal Alignment and Surface-Constrained Optimization of 3DGS for High-Fidelity Surface ReconstructionabstractMulti-view surface reconstruction is essential for accurate 3D modeling and high-quality novel view synthesis. Neural implicit methods, such as NeRF, exhibit excellent rendering but struggle with precise surface extraction. While 3D Gaussian splatting (3DGS) provides efficient explicit representations, it suffers from geometric inaccuracies due to misalignment and lack of strong surface constraints, especially in real-world scenarios. These issues stem from unordered Gaussian primitives, may tend to surface drift, redundancy, and blurred boundaries. To overcome these limitations, we propose a surface-aware Gaussian aggregation framework featuring an adaptive normal alignment loss that integrates rendered, depth-based, and monocular normals to enforce robust surface orientation supervision. Additionally, our surface-guided optimization strategy aligns Gaussian primitives precisely to surfaces by exploiting combined rendered and predicted geometric information. Extensive experiments demonstrate our approach achieves state-of-the-art surface reconstruction accuracy alongside superior novel view synthesis, with ablation studies confirming the efficacy of our contributions. Jianyu Ding, Xiaonan Mao, Jie Yang 0002, Wei Liu 0044 |
IEEE Signal Process. Lett. | 6 |
| 2025 | 2M3DF: Advancing 3D Industrial Defect Detection With Multi-Perspective Multimodal Fusion NetworkabstractIn the context of Industrial Anomaly Detection (IAD), ensuring the quality of manufactured products is critical. Traditional 2D based methods often fail to capture anomalies present in complex 3D shapes. For effective anomaly detection in 3D shapes, it is essential to incorporate global semantic context, local geometric structure, and color information of the object. To fully leverage these features, we propose a network named 2M3DF, that leverages knowledge from multi-view RGB images and corresponding point cloud information for enhanced anomaly detection performance. Our model initially employs pre-trained feature extractors that generate local features from multi-view RGB images and corresponding point clouds. The novel inter-modality feature representation and fusion module first adapts these inter-modality features and then effectively aligns and aggregates these multimodality features on a pixel-to-point basis. To learn the normality from point-wise fused multimodal features, we fit a multivariate Gaussian distribution to model the normal feature distribution. Comprehensive experimental evaluations using the MVTec3D-AD and Eyecandies dataset validate the effectiveness of our propose model and demonstrate significant improvements in comparison to existing state-of-the-art methods. Our model achieves a 96.6% mean I-AUROC while delivering real-time results. Mujtaba Asad, Waqar Azeem, Hafiz Tayyab Mustafa, Jie Yang 0002, Wei Liu 0044 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Consistency-Guided Adaptive Alternating Training for Semi-Supervised Salient Object DetectionabstractThis paper presents a novel approach that leverages two models to integrate features from numerous unlabeled images, addressing the challenge of semi-supervised salient object detection (SSOD). Unlike conventional methods that rely on selecting high-quality pseudo labels, our method identifies the model that produces consistent predictions for original images and their color transformation versions from two models to infer reliable pseudo labels for all unlabeled images, improving the diversity of the training set. Specifically, we propose adaptive selection indicators to quantify prediction differences and guide the updates of the two models using the unlabeled set alternatively. Initially, two models used in our framework are trained on the labeled set. Once the adaptive selection indicator conditions are satisfied, one model is designated as the proxy, generating pseudo labels, while the other serves as the saliency model, which is further trained using these pseudo labels. Subsequently, the updated saliency model optimizes the proxy model’s parameters according to another adaptive selection indicator. Experimental results and ablation studies on six benchmark salient object detection datasets confirm the effectiveness and robustness of our method. Our approach achieves performance comparable to recent fully supervised methods while using only one eighth of the labeled data, demonstrating its potential for efficient and scalable SSOD. This paper is publicly available athttps://github.com/Liyuan0905/CATNet. Wei Liu 0044, Hua Wang 0002, Sang-Woon Jeon, Yunliang Jiang, Zhonglong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SRS: Siamese Reconstruction-Segmentation Network Based on Dynamic-Parameter ConvolutionabstractDynamic convolution demonstrates outstanding representation capabilities, which are crucial for natural image segmentation. However, it fails when applied to medical image segmentation (MIS) and infrared small target segmentation (IRSTS) due to limited data and limited fitting capacity. In this paper, we propose a new type of dynamic convolution called dynamic parameter convolution (DPConv) which shows superior fitting capacity, and it can efficiently leverage features from deep layers of encoder in reconstruction tasks to generate DPConv kernels that adapt to input variations. Moreover, we observe that DPConv, built upon deep features derived from reconstruction tasks, significantly enhances downstream segmentation performance. We refer to the segmentation network integrated with DPConv generated from reconstruction network as the siamese reconstruction-segmentation network (SRS). We conduct extensive experiments on seven datasets including five medical datasets and two infrared datasets, and the experimental results demonstrate that our method can show superior performance over several recently proposed methods. Furthermore, the zero-shot segmentation under unseen modality demonstrates the generalization of DPConv. The code is available at: https://github.com/fidshu/SRSNet. Bingkun Nian, Fenghe Tang, Jianrui Ding, Jie Yang 0002, Zhonglong Zheng, Shaohua Kevin Zhou, Wei Liu 0044 |
IEEE Trans. Image Process. | 7 |
| 2025 | VB-KGN: Variational Bayesian Kernel Generation Networks for Motion Image DeblurringabstractMotion blur estimation is a critical and fundamental task in scene analysis and image restoration. While most state-of-the-art deep learning-based methods for single-image motion image deblurring focus on constructing deep networks or developing training strategies, the characterization of motion blur has received less attention. In this paper, we innovatively propose a non-parametric Variational Bayesian Kernel Generation Network (VB-KGN) for characterizing motion blur in a single image. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of motion blur images in a data-driven manner. The qualitative and quantitative evaluations of our experimental results demonstrate that our proposed model can generate highly accurate motion blur kernels, significantly improving motion image deblurring performance and substantially reducing the need for extensive training sample preprocessing for deblurring tasks. Ying Fu 0003, Xiaojie Li 0001, Xin Wang 0045, Xi Wu 0004, Shu Hu 0001, Siwei Lyu, Wei Liu 0044 |
IEEE Trans. Multim. | 9 |
| 2024 | SoftCLIP: Softer Cross-Modal Alignment Makes CLIP StrongerabstractDuring the preceding biennium, vision-language pre-training has achieved noteworthy success on several downstream tasks. Nevertheless, acquiring high-quality image-text pairs, where the pairs are entirely exclusive of each other, remains a challenging task, and noise exists in the commonly used datasets. To address this issue, we propose SoftCLIP, a novel approach that relaxes the strict one-to-one constraint and achieves a soft cross-modal alignment by introducing a softened target, which is generated from the fine-grained intra-modal self-similarity. The intra-modal guidance is indicative to enable two pairs have some local similarities and model many-to-many relationships between the two modalities. Besides, since the positive still dominates in the softened target distribution, we disentangle the negatives in the distribution to further boost the relation alignment with the negatives in the cross-modal learning. Extensive experiments demonstrate the effectiveness of SoftCLIP. In particular, on ImageNet zero-shot classification task, using CC3M/CC12M as pre-training dataset, SoftCLIP brings a top-1 accuracy improvement of 6.8%/7.2% over the CLIP baseline. Jinfeng Liu 0007, Enwei Zhang, Ke Li 0015, Jie Yang 0002, Wei Liu 0044, Xing Sun 0001 |
AAAI | 8 |
| 2024 | Fast Global Image Smoothing via Quasi Weighted Least Squares
Wei Liu 0044, Hongxing Qin, Xiaolin Huang, Jie Yang 0002, Michael Kwok-Po Ng |
Int. J. Comput. Vis. | 1 |
| 2024 | Image Super-Resolution via Efficient Transformer Embedding Frequency Decomposition With RestartabstractRecently, transformer-based backbones show superior performance over the convolutional counterparts in computer vision. Due to quadratic complexity with respect to the token number in global attention, local attention is always adopted in low-level image processing with linear complexity. However, the limited receptive field is harmful to the performance. In this paper, motivated by Octave convolution, we propose a transformer-based single image super-resolution (SISR) model, which explicitly embeds dynamic frequency decomposition into the standard local transformer. All the frequency components are continuously updated and re-assigned via intra-scale attention and inter-scale interaction, respectively. Specifically, the attention in low resolution is enough for low-frequency features, which not only increases the receptive field, but also decreases the complexity. Compared with the standard local transformer, the proposed FDRTran layer simultaneously decreases FLOPs and parameters. By contrast, Octave convolution only decreases FLOPs of the standard convolution, but keeps the parameter number unchanged. In addition, the restart mechanism is proposed for every a few frequency updates, which first fuses the low and high frequency, then decomposes the features again. In this way, the features can be decomposed in multiple viewpoints by learnable parameters, which avoids the risk of early saturation for frequency representation. Furthermore, based on the FDRTran layer with restart mechanism, the proposed FDRNet is the first transformer backbone for SISR which discusses the Octave design. Sufficient experiments show our model reaches state-of-the-art performance on 6 synthetic and real datasets. The code and the models are available at https://github.com/catnip1029/FDRNet. Yifan Zuo 0001, Wenhao Yao, Yuming Fang 0001, Wei Liu 0044, Yuxin Peng 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | FASTC: A Fast Attentional Framework for Semantic Traversability Classification Using Point CloudabstractProducing traversability maps and understanding the surroundings are crucial prerequisites for autonomous navigation. In this paper, we address the problem of traversability assessment using point clouds. We propose a novel pillar feature extraction module that utilizes PointNet to capture features from point clouds organized in vertical volume and a 2D encoder-decoder structure to conduct traversability classification instead of the widely used 3D convolutions. This results in less computational cost while even better performance is achieved at the same time. We then propose a new spatio-temporal attention module to fuse multi-frame information, which can properly handle the varying density problem of LIDAR point clouds, and this makes our module able to assess distant areas more accurately. Comprehensive experimental results on augmented Semantic KITTI and RELLIS-3D datasets show that our method is able to achieve superior performance over existing approaches both quantitatively and quantitatively. Our code is publicly available at https://github.com/chenyirui/FASTC. Yirui Chen, Pengjin Wei, Zhenhuan Liu, Bingchao Wang, Jie Yang 0002, Wei Liu 0044 |
ECAI | 6 |
| 2023 | Recurrent Multi-scale Transformer for High-Resolution Salient Object DetectionabstractSalient Object Detection (SOD) aims to identify and segment the most conspicuous objects in an image or video. As an important pre-processing step, it has many potential applications in multimedia and vision tasks. With the advance of imaging devices, SOD with high-resolution images is of great demand, recently. However, traditional SOD methods are largely limited to low-resolution images, making them difficult to adapt to the development of High-Resolution SOD (HRSOD). Although some HRSOD methods emerge, there are no large enough datasets for training and evaluating. Besides, current HRSOD methods generally produce incomplete object regions and irregular object boundaries. To address above issues, in this work, we first propose a new HRS10K dataset, which contains 10,500 high-quality annotated images at 2K-8K resolution. As far as we know, it is the largest dataset for the HRSOD task, which will significantly help future works in training and evaluating models. Furthermore, to improve the HRSOD performance, we propose a novel Recurrent Multi-scale Transformer (RMFormer), which recurrently utilizes shared Transformers and multi-scale refinement architectures. Thus, high-resolution saliency maps can be generated with the guidance of lower-resolution predictions. Extensive experiments on both high-resolution and low-resolution benchmarks show the effectiveness and superiority of the proposed framework. The source code and dataset are released at: https://github.com/DrowsyMon/RMFormer. Xinhao Deng 0002, Wei Liu 0044, Huchuan Lu |
ACM Multimedia | 3 |
| 2022 | CROON: Automatic Multi-LiDAR Calibration and Refinement Method in Road SceneabstractSensor-based environmental perception is a crucial part of the autonomous driving system. In order to get an excellent perception of the surrounding environment, an intelligent system would configure multiple LiDARs (3D Light Detection and Ranging) to cover the distant and near space of the car. The precision of perception relies on the quality of sensor calibration. This research aims at developing an accurate, automatic, and robust calibration strategy for multiple LiDAR systems in the general road scene. We thus propose CROON (automatic multi-LiDAR Calibration and Refinement methOd in rOad sceNe), a two-stage method including rough and refinement calibration. The first stage can calibrate the sensor from an arbitrary initial pose, and the second stage is able to precisely calibrate the sensor iteratively. Specifically, CROON utilize the nature characteristics of road scene so that it is independent and easy to apply in large-scale conditions. Experimental results on real-world and simulated data sets demonstrate the reliability and accuracy of our method. All the related data sets and codes are open-sourced on the Github website https://github.com/OpenCalib/LiDAR2LiDAR. Pengjin Wei, Guohang Yan, Yikang Li 0002, Kun Fang 0004, Xinyu Cai, Jie Yang 0002, Wei Liu 0044 |
IROS | 7 |
| 2022 | A Generalized Framework for Edge-Preserving and Structure-Preserving Image SmoothingabstractImage smoothing is a fundamental procedure in applications of both computer vision and graphics. The required smoothing properties can be different or even contradictive among different tasks. Nevertheless, the inherent smoothing nature of one smoothing operator is usually fixed and thus cannot meet the various requirements of different applications. In this paper, we first introduce the truncated Huber penalty function which shows strong flexibility under different parameter settings. A generalized framework is then proposed with the introduced truncated Huber penalty function. When combined with its strong flexibility, our framework is able to achieve diverse smoothing natures where contradictive smoothing behaviors can even be achieved. It can also yield the smoothing behavior that can seldom be achieved by previous methods, and superior performance is thus achieved in challenging cases. These together enable our framework capable of a range of applications and able to outperform the state-of-the-art approaches in several tasks. In addition, an efficient numerical solution is provided and its convergence is theoretically guaranteed even the optimization framework is non-convex and non-smooth. A simple yet effective approach is further proposed to reduce the computational cost of our method while maintaining its performance. The effectiveness and superior performance of our approach are validated through comprehensive experiments in a range of applications. Wei Liu 0044, Yinjie Lei, Xiaolin Huang, Jie Yang 0002, Michael Kwok-Po Ng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Unsupervised Image Restoration With Quality-Task-Perception LossabstractImage restoration includes various kinds of tasks, such as image denoising, image deraining and low-light image enhancement, etc. Due to the domain shift problem of current supervised methods, researchers tend to adopt unsupervised image restoration methods. However, fake color or blur image, insufficient restoration and missing semantic information are three common problems when utilizing these methods. In this paper, we propose a new hybrid loss named Quality-Task-Perception (QTP) to deal with these three problems simultaneously. Specifically, this hybrid loss includes three components: quality, task and perception. The quality part overcomes the fake color or blur image problem by enforcing image quality scores of the restored images and those of the unpaired clean images to be similar. For the task part, we tackle the insufficient restoration problem by proposing to apply a task probability network to convert the unsupervised image restoration into a supervised classification problem, and this task probability network is learned from our proposed pipeline. The perception part handles the missing semantic information by restricting the multi-scale phase consistency between the degraded image and its restored version. Comprehensive experiments on both supervised and unsupervised datasets in three image restoration tasks demonstrate the superiority of our proposed approach. Wei Xu 0050, Haoming Guo, Xiaolin Huang, Wei Liu 0044 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Image deraining with Adversarial Residual Refinement Network
Wei Xu 0050, Song Qiu, Kunyao Huang, Wei Liu 0044, Junzhe Zuo, Haoming Guo |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Looking for the Detail and Context Devils: High-Resolution Salient Object DetectionabstractIn recent years, Salient Object Detection (SOD) has shown great success with the achievements of large-scale benchmarks and deep learning techniques. However, existing SOD methods mainly focus on natural images with low-resolutions, e.g., 400×400 or less. This drawback hinders them for advanced practical applications, which need high-resolution, detail-aware results. Besides, lacking of the boundary detail and semantic context of salient objects is also a key concern for accurate SOD. To address these issues, in this work we focus on the High-Resolution Salient Object Detection (HRSOD) task. Technically, we propose the first end-to-end learnable framework, named Dual ReFinement Network (DRFNet), for fully automatic HRSOD. More specifically, the proposed DRFNet consists of a shared feature extractor and two effective refinement heads. By decoupling the detail and context information, one refinement head adopts a global-aware feature pyramid. Without increasing too much computational burden, it can boost the spatial detail information, which narrows the gap between high-level semantics and low-level details. In parallel, the other refinement head adopts hybrid dilated convolutional blocks and group-wise upsamplings, which are very efficient in extracting contextual information. Based on the dual refinements, our approach can enlarge receptive fields and obtain more discriminative features from high-resolution images. Experimental results on high-resolution benchmarks (the public DUT-HRSOD and the proposed DAVIS-SOD) demonstrate that our method is not only efficient but also performs more accurate than other state-of-the-arts. Besides, our method generalizes well on typical low-resolution benchmarks. Wei Liu 0044, Yi Zeng 0006, Yinjie Lei, Huchuan Lu |
IEEE Trans. Image Process. | 2 |
| 2021 | Semantic Scene Labeling via Deep Nested Level SetabstractSemantic scene labeling plays a very important role in intelligent transportation tasks, such as autonomous driving and advanced driver assistance. Recently, thanks to the advances of deep learning, significant improvements have been achieved for this pixel-wise labeling task. Although effective, current methods lack of explicitly modeling the boundary of objects, resulting in inaccurate labeling results. Meanwhile, traditional level set based methods perform better to capture the evolution of boundaries. However, they are sensitive to the model initialization. To address these issues, in this work we propose a novel deep learning framework, named deep nested level set (DNLS) for boundary-aware semantic scene labeling. Different from previous works, our proposed framework explicitly takes deep learned features and object boundary information into account. More specifically, our proposed framework first predicts semantic probability maps and boundary locations of objects using a bifurcated fully convolutional network (BFCN). Then, these probability maps are seamlessly integrated into a nested level set function for accurate scene labeling. As a result, our approach can automatically initialize the nested level set function, and the whole framework can be trained in an end-to-end manner, providing a new solution for accurate semantic scene parsing. Extensive experiments on public CamVid and Cityscapes datasets demonstrate that our proposed framework produces high-quality predictions with clear object boundaries and spatial consistency. Wei Liu 0044, Yinjie Lei, Huchuan Lu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | A Generalized Framework for Edge-Preserving and Structure-Preserving Image SmoothingabstractImage smoothing is a fundamental procedure in applications of both computer vision and graphics. The required smoothing properties can be different or even contradictive among different tasks. Nevertheless, the inherent smoothing nature of one smoothing operator is usually fixed and thus cannot meet the various requirements of different applications. In this paper, a non-convex non-smooth optimization framework is proposed to achieve diverse smoothing natures where even contradictive smoothing behaviors can be achieved. To this end, we first introduce the truncated Huber penalty function which has seldom been used in image smoothing. A robust framework is then proposed. When combined with the strong flexibility of the truncated Huber penalty function, our framework is capable of a range of applications and can outperform the state-of-the-art approaches in several tasks. In addition, an efficient numerical solution is provided and its convergence is theoretically guaranteed even the optimization framework is non-convex and non-smooth. The effectiveness and superior performance of our approach are validated through comprehensive experimental results in a range of applications. Wei Liu 0044, Yinjie Lei, Xiaolin Huang, Jie Yang 0002, Ian D. Reid 0001 |
AAAI | 1 |
| 2020 | Blur kernel estimation of noisy-blurred image via dynamic structure prior
Xueling Chen, Yu Zhu 0004, Wei Liu 0044, Jinqiu Sun, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2020 | Non-rigid object tracking via deep multi-scale spatial-temporal discriminative saliency maps
Wei Liu 0044, Dong Wang 0004, Yinjie Lei, Hongyu Wang 0001, Huchuan Lu |
Pattern Recognit. | 2 |
| 2020 | Embedding Bilateral Filter in Least Squares for Efficient Edge-Preserving Image SmoothingabstractEdge-preserving smoothing is a fundamental procedure for many computer vision and graphic applications. This can be achieved with either local methods or global methods. In most cases, global methods can yield superior performance over the local ones. However, local methods usually run much faster than the global ones. In this paper, we propose a new global method that embeds the bilateral filter (BLF) in the least squares (LS) model for efficient edge-preserving smoothing. The proposed method can show comparable performance with the state-of-the-art global method. Meanwhile, since the proposed method can take advantages of the efficiency of the BLF and the LS model, it runs much faster. In addition, we show the flexibility of our method which can be easily extended by replacing the BLF with its variants. They can be further modified to handle more applications. We validate the effectiveness and efficiency of the proposed method through comprehensive experiments in a range of applications. Wei Liu 0044, Chunhua Shen, Xiaolin Huang, Jie Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Deep Multiphase Level Set for Scene ParsingabstractRecently, Fully Convolutional Network (FCN) seems to be the go-to architecture for image segmentation, including semantic scene parsing. However, it is difficult for a generic FCN to predict semantic labels around the object boundaries, thus FCN-based methods usually produce parsing results with inaccurate boundaries. Meanwhile, many works have demonstrate that level set based active contours are superior to the boundary estimation in sub-pixel accuracy. However, they are quite sensitive to initial settings. To address these limitations, in this paper we propose a novel Deep Multiphase Level Set (DMLS) method for semantic scene parsing, which efficiently incorporates multiphase level sets into deep neural networks. The proposed method consists of three modules, i.e., recurrent FCNs, adaptive multiphase level set, and deeply supervised learning. More specifically, recurrent FCNs learn multi-level representations of input images with different contexts. Adaptive multiphase level set drives the discriminative contour for each semantic class, which makes use of the advantages of both global and local information. In each time-step of the recurrent FCNs, deeply supervised learning is incorporated for model training. Extensive experiments on three public benchmarks have shown that our proposed method achieves new state-of-the-art performances. The source codes will be released at https://github.com/Pchank/DMLS-for-SSP. Wei Liu 0044, Yinjie Lei, Hongyu Wang 0001, Huchuan Lu |
IEEE Trans. Image Process. | 2 |
| 2020 | RAPNet: Residual Atrous Pyramid Network for Importance-Aware Street Scene ParsingabstractStreet Scene Parsing (SSP) is a fundamental and important step for autonomous driving and traffic scene understanding. Recently, Fully Convolutional Network (FCN) based methods have delivered expressive performances with the help of large-scale dense-labeling datasets. However, in urban traffic environments, not all the labels contribute equally for making the control decision. Certain labels such as pedestrian, car, bicyclist, road lane or sidewalk would be more important in comparison with labels for vegetation, sky or building. Based on this fact, in this paper we propose a novel deep learning framework, named Residual Atrous Pyramid Network (RAPNet), for importance-aware SSP. More specifically, to incorporate the importance of various object classes, we propose an Importance-Aware Feature Selection (IAFS) mechanism which automatically selects the important features for label predictions. The IAFS can operate in each convolutional block, and the semantic features with different importance are captured in different channels so that they are automatically assigned with corresponding weights. To enhance the labeling coherence, we also propose a Residual Atrous Spatial Pyramid (RASP) module to sequentially aggregate global-to-local context information in a residual refinement manner. Extensive experiments on two public benchmarks have shown that our approach achieves new state-of-the-art performances, and can consistently obtain more accurate results on the semantic classes with high importance levels. Wei Liu 0044, Yinjie Lei, Hongyu Wang 0001, Huchuan Lu |
IEEE Trans. Image Process. | 2 |
| 2020 | Real-time Image Smoothing via Iterative Least SquaresabstractEdge-preserving image smoothing is a fundamental procedure for many computer vision and graphic applications. There is a tradeoff between the smoothing quality and the processing speed: the high smoothing quality usually requires a high computational cost, which leads to the low processing speed. In this article, we propose a new global optimization based method, named iterative least squares (ILS), for efficient edge-preserving image smoothing. Our approach can produce high-quality results but at a much lower computational cost. Comprehensive experiments demonstrate that the proposed method can produce results with little visible artifacts. Moreover, the computation of ILS can be highly parallel, which can be easily accelerated through either multi-thread computing or the GPU hardware. With the acceleration of a GTX 1080 GPU, it is able to process images of 1080p resolution (1920 × 1080) at the rate of 20fps for color images and 47fps for gray images. In addition, the ILS is flexible and can be modified to handle more applications that require different smoothing properties. Experimental results of several applications show the effectiveness and efficiency of the proposed method. The code is available at https://github.com/wliusjtu/Real-time-Image-Smoothing-via-Iterative-Least-Squares. Wei Liu 0044, Xiaolin Huang, Jie Yang 0002, Chunhua Shen, Ian D. Reid 0001 |
ACM Trans. Graph. | 1 |
| 2019 | Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene CompletionabstractSemantic Scene Completion (SSC) aims to simultaneously predict the volumetric occupancy and semantic category of a 3D scene. It helps intelligent devices to understand and interact with the surrounding scenes. Due to the high-memory requirement, current methods only produce low-resolution completion predictions, and generally lose the object details. Furthermore, they also ignore the multi-scale spatial contexts, which play a vital role for the 3D inference. To address these issues, in this work we propose a novel deep learning framework, named Cascaded Context Pyramid Network (CCPNet), to jointly infer the occupancy and semantic labels of a volumetric 3D scene from a single depth image. The proposed CCPNet improves the labeling coherence with a cascaded context pyramid. Meanwhile, based on the low-level features, it progressively restores the fine-structures of objects with Guided Residual Refinement (GRR) modules. Our proposed framework has three outstanding advantages: (1) it explicitly models the 3D spatial context for performance improvement; (2) full-resolution 3D volumes are produced with structure-preserving details; (3) light-weight models with low-memory requirements are captured with a good extensibility. Extensive experiments demonstrate that in spite of taking a single-view depth map, our proposed framework can generate high-quality SSC results, and outperforms state-of-the-art approaches on both the synthetic SUNCG and real NYU datasets. Wei Liu 0044, Yinjie Lei, Huchuan Lu, Xiaoyun Yang |
ICCV | 2 |
| 2019 | Self-paced learning with privileged information
Wei Xu 0050, Wei Liu 0044, Haoyuan Chi, Song Qiu |
Neurocomputing | 2 |
| 2019 | Video co-segmentation based on directed graph
Zhi Liu 0003, Xiaofei Zhou 0003, Wei Liu 0044, Xuemei Zou |
Multim. Tools Appl. | 4 |
| 2019 | Hyperfusion-Net: Hyper-densely reflective feature fusion for salient object detection
Wei Liu 0044, Yinjie Lei, Huchuan Lu |
Pattern Recognit. | 2 |
| 2019 | Deep gated attention networks for large-scale street-level scene segmentation
Wei Liu 0044, Hongyu Wang 0001, Yinjie Lei, Huchuan Lu |
Pattern Recognit. | 2 |
| 2019 | Salient Object Detection With Lossless Feature Reflection and Weighted Structural LossabstractSalient object detection (SOD), which aims to identify and locate the most salient pixels or regions in images, has been attracting more and more interest due to its various realworld applications. However, this vision task is quite challenging, especially under complex image scenes. Inspired by the intrinsic reflection of natural images, in this paper we propose a novel feature learning framework for large-scale salient object detection. Specifically, we design a symmetrical fully convolutional network (SFCN) to effectively learn complementary saliency features under the guidance of lossless feature reflection. The location information, together with contextual and semantic information, of salient objects are jointly utilized to supervise the proposed network for more accurate saliency predictions. In addition, to overcome the blurry boundary problem, we propose a new weighted structural loss function to ensure clear object boundaries and spatially consistent saliency. The coarse prediction results are effectively refined by these structural information for performance improvements. Extensive experiments on seven saliency detection datasets demonstrate that our approach achieves consistently superior performance and outperforms the very recent state-of-the-art methods with a large margin. Wei Liu 0044, Huchuan Lu, Chunhua Shen |
IEEE Trans. Image Process. | 2 |
| 2018 | Salient Object Detection by Lossless Feature ReflectionabstractSalient object detection, which aims to identify and locate the most salient pixels or regions in images, has been attracting more and more interest due to its various real-world applications. However, this vision task is quite challenging, especially under complex image scenes. Inspired by the intrinsic reflection of natural images, in this paper we propose a novel feature learning framework for large-scale salient object detection. Specifically, we design a symmetrical fully convolutional network (SFCN) to learn complementary saliency features under the guidance of lossless feature reflection. The location information, together with contextual and semantic information, of salient objects are jointly utilized to supervise the proposed network for more accurate saliency predictions. In addition, to overcome the blurry boundary problem, we propose a new structural loss function to learn clear object boundaries and spatially consistent saliency. The coarse prediction results are effectively refined by these structural information for performance improvements. Extensive experiments on seven saliency detection datasets demonstrate that our approach achieves consistently superior performance and outperforms the very recent state-of-the-art methods. Wei Liu 0044, Huchuan Lu, Chunhua Shen |
IJCAI | 2 |
| 2018 | Multi-modal self-paced learning for image classification
Wei Xu 0050, Wei Liu 0044, Xiaolin Huang, Jie Yang 0002, Song Qiu |
Neurocomputing | 2 |
| 2018 | Multi-task classification with sequential instances and tasks
Wei Xu 0050, Wei Liu 0044, Haoyuan Chi, Xiaolin Huang, Jie Yang 0002 |
Signal Process. Image Commun. | 2 |
| 2018 | Improving Video Saliency Detection via Localized Estimation and Spatiotemporal RefinementabstractVideo saliency detection aims to pop out the most salient regions in every frame of a video. Up to now, many efforts have been made from various aspects for video saliency detection. Unfortunately, the existing video saliency models are very likely to fail in challenging videos with complicated motions and complex scenes. Therefore, in this paper, we propose a novel framework to improve the saliency detection results generated by existing video saliency models. The proposed framework consists of three key steps including localized estimation, spatiotemporal refinement, and saliency update. Specifically, the initial saliency map of each frame in a video is first generated by using an existing saliency model. Then, by considering the temporal consistency and strong correlation among adjacent frames, the localized estimation models, which are generated by training the random forest regressor within a local temporal window, are employed to generate the temporary saliency map. Finally, by taking the appearance and motion information of salient objects into consideration, the spatiotemporal refinement step is deployed to further improve the temporary saliency map and generate the final saliency map. Furthermore, such an improved saliency map is then utilized to update the initial saliency map and provide reliable cues for saliency detection in the next frame. The experimental results on four challenging datasets demonstrate that the proposed framework is able to consistently and significantly improve the saliency detection performance of various video saliency models, thereby achieving the state-of-the-art performance. Xiaofei Zhou 0003, Zhi Liu 0003, Chen Gong 0002, Wei Liu 0044 |
IEEE Trans. Multim. | 4 |
| 2017 | Semi-Global Weighted Least Squares in Image FilteringabstractSolving the global method of Weighted Least Squares (WLS) model in image filtering is both time- and memory-consuming. In this paper, we present an alternative approximation in a time- and memory- efficient manner which is denoted as Semi-Global Weighed Least Squares (SG-WLS). Instead of solving a large linear system, we propose to iteratively solve a sequence of subsystems which are one-dimensional WLS models. Although each subsystem is one-dimensional, it can take two-dimensional neighborhood information into account due to the proposed special neighborhood construction. We show such a desirable property makes our SG-WLS achieve close performance to the original two-dimensional WLS model but with much less time and memory cost. While previous related methods mainly focus on the 4-connected/8-connected neighborhood system, our SG-WLS can handle a more general and larger neighborhood system thanks to the proposed fast solution. We show such a generalization can achieve better performance than the 4-connected/8-connected neighborhood system in some applications. Our SG-WLS is ~ 20 times faster than the WLS model. For an image of M × N, the memory cost of SG-WLS is at most at the magnitude of max{1/M,1/N} of that of the WLS model. We show the effectiveness and efficiency of our SG-WLS in a range of applications. Wei Liu 0044, Chunhua Shen, Zhi Liu 0003, Jie Yang 0002 |
ICCV | 1 |
| 2017 | Variable Bandwidth Weighting for Texture Copy Artifact Suppression in Guided Depth UpsamplingabstractIn this paper, we mathematically analyze one of the most challenging issues in color image-guided depth upsampling: the texture copy artifacts. The optimal guidance weights denoted by balanced weights are proposed to best suppress texture copy artifacts. To both suppress texture copy artifacts and preserve depth discontinuities, a new general weighting scheme called variable bandwidth weighting is proposed. The variable bandwidth weighting scheme is able to adjust guidance weights according to the local depth smoothness. A new concept called relative smoothness is proposed for measuring the local depth smoothness. Given this quantitative smoothness measurement, the proposed weighting scheme can adaptively adjust the bandwidth for calculating the guidance weights in the existing methods. As we use the computationally efficient balanced weights instead of the guidance weights of a large bandwidth, the proposed method can speed up the upsampling process for about $2\times \sim 5\times $ when compared with the original upsampling methods. Experimental results show the effectiveness and efficiency of the proposed method in suppressing texture copy artifacts, preserving the depth discontinuities and reducing the computational cost at the same time. Wei Liu 0044, Jie Yang 0002, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Robust Color Guided Depth Map RestorationabstractOne of the most challenging issues in color guided depth map restoration is the inconsistency between color edges in guidance color images and depth discontinuities on depth maps. This makes the restored depth map suffer from texture copy artifacts and blurring depth discontinuities. To handle this problem, most state-of-the-art methods design complex guidance weight based on guidance color images and heuristically make use of the bicubic interpolation of the input depth map. In this paper, we show that using bicubic interpolated depth map can blur depth discontinuities when the upsampling factor is large and the input depth map contains large holes and heavy noise. In contrast, we propose a robust optimization framework for color guided depth map restoration. By adopting a robust penalty function to model the smoothness term of our model, we show that the proposed method is robust against the inconsistency between color edges and depth discontinuities even when we use simple guidance weight. To the best of our knowledge, we are the first to solve this problem with a principled mathematical formulation rather than previous heuristic weighting schemes. The proposed robust method performs well in suppressing texture copy artifacts. Moreover, it can better preserve sharp depth discontinuities than previous heuristic weighting schemes. Through comprehensive experiments on both simulated data and real data, we show promising performance of the proposed method. Wei Liu 0044, Jie Yang 0002, Qiang Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Robust weighted least squares for guided depth upsamplingabstractIn this paper, we propose a new guided depth upsampling method denoted as Robust Weighted Least Squares (RWLS). Our work is inspired by the connection between the Weighted Least Squares (WLS) and the Auto Regressive (AR) model. By adopting a new robust penalty function to model the smoothness of the proposed model, we show that the proposed method performs much better in preserving sharp depth discontinuities than previous work. Through both mathematical analysis and experimental results, we show that our method has promising performance on handling the inconsistency between the guidance image and the depth map in both preserving sharp depth discontinuities and suppressing the texture copy artifacts. Wei Liu 0044, Jie Yang 0002, Qiang Wu 0001 |
ICIP | 1 |
| 2016 | Unsupervised Video Hashing by Exploiting Spatio-Temporal Feature
Chao Ma 0005, Yun Gu, Wei Liu 0044, Jie Yang 0002, Xiangjian He |
ICONIP (3) | 3 |
| 2015 | Upsampling the depth map with its own propertiesabstractIn this paper, we present a novel method to upsample the depth map obtained by the Time-of-Flight (ToF) camera with the guidance of the companion high resolution color image. The problem is modeled with an optimization framework where we use a novel exponential function as the error norm. By using this novel error norm, our model could take the properties of the depth map itself into account. Depth discontinuity cues are obtained not only from the color image but also the depth map itself. To further enhance the performance, we perform a data driven selection of the parameter in the model to better fit the property of the depth map. Experimental results show that our method has excellent performance in smoothing the noise, preserving sharp depth discontinuities and suppressing the texture copy effect. Wei Liu 0044, Penglin Li, Jie Yang 0002 |
ICIP | 1 |
| 2015 | An MRF-Based Depth Upsampling: Upsample the Depth Map With Its Own PropertyabstractIn this letter, we propose a novel method for upsampling the noisy low resolution depth map with the guidance of the companion color image. The problem is modeled with an Markov Random Field (MRF)-based optimization framework. The novelty relies on the smoothness term that is modeled with an exponential function as the error norm. By using this novel error norm, our method can take the property of the depth map into account. Depth discontinuity cues are not only obtained from the color image but also the depth map itself. Our method has much better performance in preserving sharp depth discontinuities and suppressing the texture copy artifacts. Experimental results show that our method outperforms state-of-art solutions in both visual quality and accuracy. Wei Liu 0044, Shaoyong Jia, Penglin Li, Jie Yang 0002, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2014 | Shape Preserving RGB-D Depth Map Restoration
Wei Liu 0044, Haoyang Xue, Yun Gu, Jie Yang 0002, Qiang Wu 0001, Zhenhong Jia |
ICONIP (3) | 1 |