VLDB 2026 Research / reviewers in the wild / expert
Weisheng Dong
dblp:03/5833
· DBLP profile ↗
167ranked-venue papers
26as first author
111since 2021 · last 2026
0000-0002-9632-985XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 120 · 22 first-author · 78 since 2021Artificial intelligence and machine learning · 64 · 6 first-author · 45 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-layer graph constraint dictionary pair learning for image classification
Guangming Shi, Weisheng Dong, Xuemei Xie |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | CoCoFR: Collaborative codebooks learning with soft matching strategy for blind face restoration
Teng Feng, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi |
Neural Networks | 6 |
| 2026 | Semantic consistency-aware pseudo-temporal framework for multimodal remote sensing image segmentation
Yuejiang Li, Weisheng Dong, Peng Wu 0015, Lichao Mou, Xin Li 0005 |
Neural Networks | 3 |
| 2026 | Visible-infrared joint image deraining for harsh rain conditions with cross-modal semantic consistency
Xin Li 0005, Chengpei Xu, Zhenyu Wang 0008, Weisheng Dong |
Pattern Recognit. | 7 |
| 2026 | Learning coherent matrixized representation in latent space for volumetric 4D generation
Qitong Yang, Mingtao Feng, Shijie Sun 0001, Weisheng Dong, Yaonan Wang 0001, Mian M. Ajmal |
Pattern Recognit. | 5 |
| 2026 | Description helps: Semantic and texture consistency constraints for SAR-to-optical translation
Fanhao Zhou, Mingtao Feng, Weisheng Dong |
Pattern Recognit. | 5 |
| 2026 | Second-Order Robust Iterative Pose Optimization for Fine-Grained Cross-View LocalizationabstractFine-grained cross-view localization seeks to estimate precise camera poses by matching ground images with GPS-tagged aerial imagery. Existing methods typically employ first-order iterative optimization to progressively update the camera pose based on cross-view feature correspondences. However, they rely on local features and neglect global and complementary contextual information, making them prone to local optima and slow convergence under large initial errors or strong disturbances. To overcome these limitations, we propose a second-order robust iterative pose estimation framework for fine-grained cross-view localization. Firstly, we devise a second-order deep iterative optimization module to capture complementary forward and backward motion cues, leading to a bidirectional correlation volume. A motion aggregator uses the volume to approximate the dynamics of second-order iterators, substantially facilitating convergence and robustness. In addition, a bidirectional motion-aware robust regularization module mitigates geometric distortions and outlier interference by leveraging bidirectional motion cues to generate fine-grained confidence maps, adaptively suppressing unreliable regions and enhancing the stability of iterative optimization and pose estimation accuracy. Extensive experiments demonstrate that the proposed framework achieves faster convergence and higher pose estimation accuracy than state-of-the-art methods, particularly under large initial errors and challenging conditions. Mingtao Feng, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Image Process. | 4 |
| 2026 | LoopExpose: An Unsupervised Framework for Arbitrary-Length Exposure CorrectionabstractExposure correction is essential for enhancing image quality under challenging lighting conditions. While supervised learning has achieved significant progress in this area, it relies heavily on large-scale labeled datasets, which are difficult to obtain in practical scenarios. To address this limitation, we propose a pseudo label-based unsupervised method called LoopExpose for arbitrary-length exposure correction. A nested loop optimization strategy is proposed to address the exposure correction problem, where the correction model and pseudo-supervised information are jointly optimized in a two-level framework. Specifically, the upper-level trains a correction model using pseudo-labels generated through multi-exposure fusion at the lower level. A feedback mechanism is introduced where corrected images are fed back into the fusion process to refine the pseudo-labels, creating a self-reinforcing learning loop. Considering the dominant role of luminance calibration in exposure correction, a Luminance Ranking Loss is introduced to leverage the relative luminance ordering across the input sequence as a self-supervised constraint. Extensive experiments on different benchmark datasets demonstrate that LoopExpose achieves superior exposure correction and fusion performance, outperforming existing state-of-the-art unsupervised methods. Code is available at https://github.com/FALALAS/LoopExpose. Zhenyu Wang 0008, Weisheng Dong |
IEEE Trans. Image Process. | 6 |
| 2026 | Learning Retinex Prior for Compressive Hyperspectral Image ReconstructionabstractImage reconstruction in coded aperture snapshot spectral compressive imaging (CASSI) aims to recover high-fidelity hyperspectral images (HSIs) from compressed 2D measurements. While deep unfolding networks have shown promising performance, the degradation induced by the CASSI degradation model often introduces global illumination discrepancies in the reconstructions, creating artifacts similar to those in low-light images. To address these challenges, we propose a novel Retinex Prior-Driven Unfolding Network (RPDUN), which unfolds the optimization incorporating the Retinex prior as a regularization term into a multi-stage network. This design provides global illumination adjustment for compressed measurements, effectively compensating for spatial-spectral degradation according to physical modulation and capturing intrinsic spectral characteristics. To the best of our knowledge, this is the first application of the Retinex prior in hyperspectral image reconstruction. Furthermore, to mitigate the noise in the reflectance domain, which can be amplified during decomposition, we introduce an Adaptive Token Selection Transformer (ATST). This module adaptively filters out weakly correlated tokens before the self-attention computation, effectively reducing noise and artifacts within the recovered reflectance map. Extensive experiments on both simulated and real-world datasets demonstrate that RPDUN achieves new state-of-the-art performance, significantly improving reconstruction quality while maintaining computational efficiency. The code is available at https://github.com/ZUGE0312/RPDUN. Mengzu Liu, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 3 |
| 2026 | Uncertainty-Driven Generative Prior Learning for Sparse Model-Guided Hyperspectral Image FusionabstractAs an alternative to acquiring high-resolution hyperspectral images (HR-HSI), Hyperspectral Image Fusion (HIF) aims to recover clean HR-HSIs by fusing degraded low spatial resolution hyperspectral images and high spatial resolution multispectral images. Among existing HIF approaches, model-guided HIF methods stand out by integrating physical degradation constraints with the learning capabilities of data-driven networks. However, most of them learn deep priors only from degraded-clean pairs without degradation-free knowledge, making them struggle with severe or unseen degradations. To address these issues, we propose a Vector-Quantized Prior-Guided Network (VPG-Net), an unfolding-based HIF framework enhanced by sparse representation and novel uncertainty-driven generative priors. Specifically, VPG-Net unfolds the Maximum A Posteriori (MAP) estimation with a sparse representation model into an uncertainty-aware VQ prior-guided network implementation. Within this framework, the sparse representation prior is integrated into the MAP formulation to improve noise resistance. As the core of our method, we leverage a high-quality vector-quantized (VQ) prior, which serves as a powerful degradation-free generative prior for the HIF process. We pre-train a discrete codebook and encoder on clean HR-HSIs to generate a VQ-prior representation (VQPR), which preserves complete spatial-spectral information. To effectively bridge the gap between degraded inputs and the learned degradation-free codebook, we further incorporate a novel uncertainty-driven probabilistic matching strategy that improves feature alignment and suppresses artifacts. The learned VQPR is then incorporated into the deep prior module as dynamic modulation parameters to enhance the fidelity and realism of the reconstructed results, particularly for severely degraded inputs. Extensive experiments on clean and degraded synthetic and real-world datasets demonstrate that our approach outperforms state-of-the-art HIF methods in both quantitative metrics and visual quality. Teng Feng, Zhenxuan Fang, Weisheng Dong, Xin Li 0005 |
IEEE Trans. Image Process. | 8 |
| 2025 | Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-LocalizationabstractVisual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing. This makes it challenging to automatically control the relative importance of the two tasks, potentially compromising the retrieval task in favor of the offset regression. Consequently, the model may encounter conflicting dominating gradients during joint training. To address this, we propose to model the semantic ambiguity during the offset regression process by integrating associated uncertainty scores, represented as 2D Gaussian distributions, to mitigate negative transfer effects within the joint tasks. We further introduce an uncertainty-aware similarity metric to enhance similarity assessment between query and reference images, accounting for their semantic ambiguities. This metric propagates uncertainty scores into the retrieval task, focusing on certain samples and learning discriminative feature embeddings, allowing the model to adaptively handle conflicting dominating gradients during joint training. Extensive experiments demonstrate that our method improves the overall performance of the joint tasks, achieving state-of-the-art results on the VIGOR and CVACT datasets. Mingtao Feng, Fenghao Tian, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
AAAI | 5 |
| 2025 | Asymmetric Hierarchical Difference-aware Interaction Network for Event-guided Motion DeblurringabstractEvent cameras are bio-inspired sensors that are capable of capturing motion information with high temporal resolution, which show potential in aiding image motion deblurring recently. Most existing methods indiscriminately handle feature fusion of two modalities with symmetric unidirectional/bidirectional interactions at different-level layers in feature encoder, while ignoring the different dependencies between cross-modal hierarchical features. To tackle these limitations, we propose a novel Asymmetric Hierarchical Difference-aware Interaction Network (AHDINet) for event-based motion deblurring, which explores the complementarity of two modalities with differential dependence modeling of cross-modal hierarchical features. Thereby, an event-assisted edge complement module is designed to leverage event modality to enhance the edge details of the image features in low-level encoder stage, and an image-assisted semantic complement module is developed to transfer contextual semantics of image features to event branch in high-level encoder stage. Benefiting from the proposed differentiated interaction mode, the respective advantages of image and event modalities are fully exploited. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance. Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
AAAI | 4 |
| 2025 | Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object DetectionabstractTiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture to enhance the extinguished regions. We for the first time exploit the regions to be enhanced from the perspective of pixel-wise amount of information. Specifically, we model the entire image pixels feature information by minimizing Information Entropy loss, generating an information map to attentively highlight weak activated regions in an unsupervised way. To effectively assist the above phase with more attention to tiny objects, we next introduce the Position Gaussian Distribution Map, explicitly modeled using a Gaussian Mixture distribution, where each Gaussian component's parameters depend on the position and size of object instance labels, serving as supervision for further feature enhancement. Taking the information map as prior knowledge guidance, we construct a multi-scale position gaussian distribution map prediction module, simultaneously modulating the information map and distribution map to focus on tiny objects during training. Extensive experiments on three public tiny object datasets demonstrate the superiority of our method over current state-of-the-art competitors. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
CVPR | 3 |
| 2025 | Parameterized Blur Kernel Prior Learning for Local Motion Deblurring
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Guangming Shi |
CVPR | 5 |
| 2025 | Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring ClassesabstractRecent approaches, such as data augmentation, adversarial training, and transfer learning, have shown potential in addressing the issue of performance degradation caused by distributional shifts. However, they typically demand careful design in terms of data or models and lack awareness of the impact of distributional shifts. In this paper, we observe that classification errors arising from distribution shifts tend to cluster near the true values, suggesting that misclassifications commonly occur in semantically similar, neighboring categories. Furthermore, robust advanced vision foundation models maintain larger inter-class distances while preserving semantic consistency, making them less vulnerable to such shifts. Building on these findings, we propose a new method called GFN (Gain From Neighbors), which uses gradient priors from neighboring classes to perturb input images and incorporates an inter-class distance-weighted loss to improve class separation. This approach encourages the model to learn more resilient features from data prone to errors, enhancing its robustness against shifts in diverse settings. In extensive experiments across various model architectures and benchmark datasets, GFN consistently demonstrated superior performance. For instance, compared to the current state-of-the-art TAPADL method, our approach achieved a higher corruption robustness of 41.4% on ImageNet-C (+2.3%), without requiring additional parameters and using only minimal data. Mingtao Feng, Weisheng Dong, Xin Li 0005, Guangming Shi |
CVPR | 5 |
| 2025 | Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generationabstract3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework designed to efficiently and controllably generate high-resolution 3D models from text prompts or images. Our key insights are three-fold: 1) Hierarchical Gaussian Mixture Model Splatting: We propose a hybrid hierarchical representation to extract fixed number of fine-grained Gaussians with multiscale details from textured object, also establish part-level representation of Gaussians primitives. 2) Mamba with adaptive tree topology: We present a diffusion mamba with tree-topology to adaptively generate Gaussians with disordered spatial structures, without the need for complex preprocessing and maintain linear complexity generation. 3) Controllable Generation: Building on the HGMM tree, we introduce a cascaded diffusion framework combining controllable implicit latent generation, which progressively generates condition-driven latents, and explicit splatting generation, which transforms latents into high-quality Gaussian primitives. Extensive experiments demonstrate the high fidelity and efficiency of our approach. Qitong Yang, Mingtao Feng, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
CVPR | 4 |
| 2025 | Bridging Task Boundaries: Remote Sensing Image-Text Retrieval via Dictionary-Driven AdaptationabstractGiven image (or text), remote sensing image-text retrieval (RSITR) aims to retrieve corresponding text (or image) within diverse remote sensing data. However, due to the complex scenes and compact distribution of targets in remote sensing data, existing methods, particularly those leveraging large models like CLIP, often generate features with high intra-modal similarity and insufficient distinctive characteristics, thus resulting in suboptimal retrieval performance. To address these issues, we pro- pose a novel dictionary-based RSITR method that jointly models image and text feature estimation. Specifically, by incorporating a general dictionary and the corresponding sparse coefficients, our method more effectively captures the correlations between the learned features. Furthermore, we introduce adaptive weighted metric learning on a sample-by-sample basis to encourage the model to focus on more challenging negative samples, promoting fine-grained feature alignment. The extensive experiments on the RSICD and RSITMD datasets demonstrate the effectiveness of our method, demonstrating significant improvements in retrieval performance. Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005 |
ICASSP | 4 |
| 2025 | Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization
Mingtao Feng, Longlong Mei, Jianqiao Luo, Fenghao Tian, Jie Feng 0003, Weisheng Dong, Yaonan Wang 0001 |
ICCV | 7 |
| 2025 | PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote SensingabstractRemote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patterncom in RSCIR has significant flaws, including the lack of query text sentences and paired triplets, thus making it unable to evaluate the latest methods. To address this, we propose the Zero-Shot Query Text Generator (ZS-QTG) that can generate full query text sentences based on attributes, and then, by capitalizing on ZS-QTG, we develop the PatternCIR benchmark. Pattern CIR rectifies Patterncom’s deficiencies and enables the evaluation of existing methods. Additionally, we explore zero-shot composed image retrieval methods that do not rely on massive pre-collected triplets for training. Existing methods use only the text during retrieval, performing poorly in RSCIR. To improve this, we propose Text-image Sequential Training of Composed Image Retrieval (TisCIR). TisCIR undergoes sequential training of multiple self-masking projection and fine-grained image attention modules, which endows it with the capacity to filter out conflicting information between the image and text, enhancing the retrieval by utilizing both modalities in harmony. TisCIR outperforms existing methods by 12.40% to 62.03% on PatternCIR, achieving state-of-the-art performance in RSCIR. The data and code are available here. Zhechun Liang, Shiwen Xue, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi |
IJCAI | 6 |
| 2025 | High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
Jinghao Bian, Jingyang Hou, Jingliang Hu, Yilei Shi, Weisheng Dong, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (14) | 6 |
| 2025 | Pushing the Limit of Binarized Neural Network for Image Super Resolution with Smooth Information TransmissionabstractLightweight models are currently the focal point in image super-resolution (ISR) research, of which the application on resource-limited devices is constrained by heavy computational requirements. As an efficient approach to enhance the inference efficiency of deep learning models, low-bit quantization has garnered significant interest. In this paper, we emphasize that low-bit ISR is not merely a parody of its full-precision version and explore binary quantization in ISR from the perspective of information transmission, pushing the limits of binarized ISR. Specifically, we propose a Maximum Entropy Routing (MER) mechanism to dynamically control activation distribution, maximizing the information entropy of binarized feature maps. Additionally, a Learnable Deviation Compensation (LDC) and an Adaptive Step-size Estimation (ASE) are introduced to reduce information loss during the forward and backward passes, respectively. By enabling smoother information transmission through more flexible binarized activation representations and more precise gradient estimation, the performance gap between binarized and full-precision models is narrowed to less than 0.3 dB. Extensive experiments demonstrate that our proposed binarization method achieves state-of-the-art results in Peak Signal-to-Noise Ratio (PSNR) across all popular benchmarks. Weimin Cheng, Zhenyu Wang 0008, Weisheng Dong |
ACM Multimedia | 5 |
| 2025 | Beyond Visual Quality: Fidelity-Oriented Diffusion Model for Real-world Image Super-ResolutionabstractAlthough existing diffusion-based image super-resolution methods have achieved remarkable visual quality, they often struggle with fidelity issues, particularly in preserving consistency with the original input image. This issue arises because using low-quality images as conditional inputs introduces substantial errors in the diffusion backward denoising process, making the restored features deviate from target features and thus degrade image fidelity. To improve the accuracy of noise estimation, we propose a dual-memory module to reinforce the input low-quality conditional features, which consists of a pre-trained high-quality memory bank to enrich the structural information and a degradation memory to remove the degradation components. Furthermore, we develop an uncertainty-aware noise estimation framework, utilizing an extra branch in the denoising network to predict pixel-wise uncertainty values, thus dynamically adjust the optimization weights for high-uncertainty regions. This adaptive strategy effectively improves the accuracy of noise estimation in challenging reconstruction areas. Experimental results demonstrate that our method significantly enhances the fidelity while preserving high visual quality of diffusion-based super-resolution, improving the reliability of diffusion applications. Zhenxuan Fang, Shuaibo Wang, Weisheng Dong, Xin Li 0005, Guangming Shi |
ACM Multimedia | 3 |
| 2025 | Exploring Global Correlations via Polarity Memory for Multispectral DemosaicingabstractMultispectral image demosaicing aims to reconstruct full band multispectral images from a compressed spectral mosaic images. Although existing learning-based methods have made progress in multispectral image demosaicing, there still exist intrinsic performance bottlenecks due to the heavy undersampling according to mosaic pattern. To address this issue, we propose Polarity memory network with quant attention to establish global correlation, thus reconstructing high-quality multispectral images from compressed spectral mosaic images. Our proposed Polarity memory network adaptively encapsulates reconstruction-oriented representations, then amplifies relevant ones and reducing noise from irrelevant ones in a polarity-aware manner to better cater to the enhancement of different spectral information with linear computational complexity. Moreover, considering existing methods' inability to adequately compensate for long distance interactions in reconstruction, we introduce a quant attention paradigm that categorize tokens into semantic-aware groups using an efficient quant operation for attention computation. Experimental results show our method achieves state-of-the-art performance on various simulation datasets and better vision results on real-world datasets. Mengzu Liu, Xin Li 0005, Weisheng Dong |
ACM Multimedia | 7 |
| 2025 | Generalizing to New Area: Self-Distillation Curriculum Learning for Fine-Grained Cross View LocalizationabstractFine-grained cross-view localization seeks to predict ground-level camera positions within GPS-tagged aerial images by matching ground and aerial views. Existing methods often rely on large-scale ground truth annotations from specific regions, but performance degrades due to domain shifts when models trained in one area are applied to another. However, collecting region-specific annotations for each area is costly or infeasible. To address this, we propose a self-distillation curriculum learning framework that generalizes pretrained localization models to unseen new areas. Our approach introduces a Dirichlet-based quality assessment strategy to evaluate teacher-generated pseudo labels, where high uncertainty signals noisy predictions and low uncertainty indicates clean samples. This uncertainty is used to guide an easy-to-hard curriculum learning strategy, where easy samples are prioritized initially, and more challenging samples are progressively incorporated, enabling effective student training. Furthermore, we develop a joint optimization scheme that updates both the student model and pseudo labels, applying adaptive label smoothing to mitigate label noises and taking full advantage of new area data. Extensive experimental results on the VIGOR and KITTI benchmarks demonstrate that our method outperforms state-of-the-art approaches in new area localization, achieving superior accuracy without additional supervision. Fenghao Tian, Mingtao Feng, Jianqiao Luo, Longlong Mei, Weisheng Dong, Yaonan Wang 0001 |
ACM Multimedia | 7 |
| 2025 | Towards Syn-to-Real IQA: A Novel Perspective on Reshaping Synthetic Data DistributionsabstractBlind Image Quality Assessment (BIQA) has advanced significantly through deep learning, but the scarcity of large-scale labeled datasets remains a challenge. While synthetic data offers a promising solution, models trained on existing synthetic datasets often show limited generalization ability. In this work, we make a key observation that representations learned from synthetic datasets often exhibit a discrete and clustered pattern that hinders regression performance: features of high-quality images cluster around reference images, while those of low-quality images cluster based on distortion types. Our analysis reveals that this issue stems from the distribution of synthetic data rather than model architecture. Consequently, we introduce a novel framework SynDR-IQA, which reshapes synthetic data distribution to enhance BIQA generalization. Based on theoretical derivations of sample diversity and redundancy's impact on generalization error, SynDR-IQA employs two strategies: distribution-aware diverse content upsampling, which enhances visual diversity while preserving content distribution, and density-aware redundant cluster downsampling, which balances samples by reducing the density of densely clustered areas. Extensive experiments across three cross-dataset settings (synthetic-to-authentic, synthetic-to-algorithmic, and synthetic-to-synthetic) demonstrate the effectiveness of our method. The code is available at https://github.com/Li-aobo/SynDR-IQA. Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong |
NeurIPS | 5 |
| 2025 | Disambiguating Holistic Language for 3D Visual Grounding via Neural-Symbolic Reasoning
Yunze Wu, Yufan Zhu, Mingtao Feng, Weisheng Dong |
PRCV (17) | 6 |
| 2025 | Compressing Vision Transformer from the View of Model Property in Frequency Domain
Zhenyu Wang 0008, Xuemei Xie, Hao Luo 0004, Weisheng Dong, Yongxu Liu 0001, Fan Wang 0019, Guangming Shi |
Int. J. Comput. Vis. | 5 |
| 2025 | Searching efficient network with lightweight optimization for image super-resolution
Ruina Shen, Zhangheng Peng, Weisheng Dong, Guangming Shi |
Neurocomputing | 5 |
| 2025 | Fast Window-Based Event Denoising With Spatiotemporal Correlation EnhancementabstractPrevious deep learning-based event denoising methods mostly suffer from poor interpretability and difficulty in real-time processing due to their complex architecture designs. In this paper, we propose window-based event denoising, which simultaneously deals with a stack of events while existing element-based denoising focuses on one event each time. Besides, we give the theoretical analysis based on probability distributions in both temporal and spatial domains to improve interpretability. In temporal domain, we use timestamp deviations between processing events and central event to judge the temporal correlation and filter out temporal-irrelevant events. In spatial domain, we choose maximum a posteriori (MAP) to discriminate real-world event and noise and use the learned convolutional sparse coding to optimize the objective function. Based on the theoretical analysis, we build Temporal Window (TW) module and Soft Spatial Feature Embedding (SSFE) module to process temporal and spatial information separately, and construct a novel multi-scale window-based event denoising network, named WedNet. The high denoising accuracy and fast running speed of our WedNet enables us to achieve real-time denoising in complex scenes. Extensive experimental results verify the effectiveness and robustness of our WedNet. Our algorithm can remove event noise effectively and efficiently and improve the performance of downstream tasks. Huachen Fang, Jinjian Wu, Qibin Hou, Weisheng Dong, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Hyperrectangle Embedding for Debiased 3D Scene Graph Prediction From RGB Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment and is regarded as a prerequisite for long-term autonomous robotic operations. A practical research problem here is to predict the 3D scene graph from sequentially captured data. However, existing methods neglect the polysemy of semantic roles that coarse feature vectors are insufficient to represent entities in different relationship semantics. This extremely limits their capability to predict relationships. We propose an approach to tackle the aforementioned challenge by introducing a novel representation, the hyperrectangle embedding, which represents entity using distinctive geometry for more effective scene understanding, rather than learning within vector-based feature with blindly increasing dimensions. By incorporating an entity within two affine-transformed embeddings, each representing either the subject or object and characterized by separate learnable transformations, we achieve the polysemy of semantic roles. The intersections of affine-transformed hyperrectangle embeddings represent the bidirectional relationship between two entities. We identify bias and reliability as two challenges impeding the model learning process. In response to the bias, that arises from long-tailed distributions in the data, we propose a history-guided debiasing strategy that utilizes a confusion history block comprised of previous hyperrectangle embeddings. This strategy mitigates inherent biases by extracting pertinent information and facilitating knowledge transfer from dominant categories to rare ones. To enhance the reliability of predictions, we introduce predictive uncertainty into the 3D scene graph prediction task. We develop a post-hoc reliability enhancement strategy to identify potentially unreliable predictions and subsequently enhance the model's predictive accuracy. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming existing state-of-the-art. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Growing-before-pruning: A progressive neural architecture search strategy via group sparsity and deterministic annealingabstractNetwork pruning is a widely studied technique of obtaining compact representations from over-parameterized deep convolutional neural networks . Existing pruning methods are based on finding an optimal combination of pruned filters in the fixed search space . However, the optimality of those methods is often questionable due to limited search space and pruning choices - e.g., the difficulty with removing the entire layer and the risk of unexpected performance degradation . Inspired by the exploration vs. exploitation trade-off in reinforcement learning, we propose to reconstruct the filter space without increasing the model capacity and prune them by exploiting group sparsity . Our approach challenges the conventional wisdom by advocating the strategy of Growing-before-Pruning (GbP), which allows us to explore more space before exploiting the power of architecture search. Meanwhile, to achieve more efficient pruning, we propose to measure the importance of filters by global group sparsity , which extends the existing Gaussian scale mixture model. Such global characterization of sparsity in the filter space leads to a novel deterministic annealing strategy for progressively pruning the filters. We have evaluated our method on several popular datasets and network architectures. Our extensive experiment results have shown that the proposed method advances the current state-of-the-art. Xiaotong Lu, Weisheng Dong, Zhenxuan Fang, Jie Lin 0008, Xin Li 0005, Guangming Shi |
Pattern Recognit. | 2 |
| 2025 | History-Enhanced 3D Scene Graph Reasoning From RGB-D Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment, and is considered a prerequisite for long-term autonomous robotic operations. However, building rich representations from RGB-D sequences remains a challenging problem. Existing methods ignore the semantic gap between linguistic and geometric feature spaces or neglect the importance of historical context in incrementally captured data. This limits the learning of visual-textual correspondence and the capability of relationship prediction. To address these problems, we propose a history-enhanced 3D scene graph reasoning framework that incrementally builds a consistent 3D semantic scene graph from an RGB-D image sequence. Specifically, we first introduce a cross-domain unified feature representation module to describe the object instances and their relationships distinctly. Next, we build a one-hot candidate matrix-enabled recurrent mechanism to reason the 3D scene graph, combining the perceived global and local history information. Finally, we design history-aware supervised semantics contrastive learning to optimize the scene-specific global history features. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be available athttps://github.com/cbyan1003/HE-3DSGR. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Towards Explainable Image Aesthetics Assessment With Attribute-Oriented Critiques GenerationabstractCompared with the unimodal image aesthetics assessment (IAA), multimodal IAA has demonstrated superior performance. This indicates that the critiques could provide rich aesthetics-aware semantic information, which also enhance the explainability of IAA models. However, images are not always accompanied with critiques in real-world situation, rendering multimodal IAA inapplicable in most cases. Therefore, it would be interesting to investigate whether we can generate aesthetic critiques to facilitate image aesthetic representation learning and enhance model explainability. Motivated by these facts, this paper presents an attribute-oriented Critiques Generation framework for explainable IAA, dubbed CG-IAA, which consists of three major components, i.e., Vision-Language Aesthetic Pretraining (VLAP), Multi-Attribute Experts Learning (MAEL) and Multimodal Aesthetics Prediction (MAP). Specifically, the vanilla CLIP is first finetuned on a multimodal IAA database. Considering that the aesthetic critiques typically consist of multiple attributes, a new multimodal IAA database which contains over 1 million critiques with up to four aesthetic attributes is constructed with the language model-based knowledge transfer. Then, CLIP-based multi-attribute experts are trained based on this database. Finally, the pretrained experts are utilized to generate aesthetic critiques for assisting unimodal image aesthetics prediction. Extensive experiments have been done on four popular IAA databases, and the results demonstrate the advantage of CG-IAA over the state-of-the-arts. Furthermore, CG-IAA features better explainability and generalization with the assistance of generated critiques. The source code is available athttps://github.com/sxfly99/CG-IAA. Leida Li, Xiangfei Sheng, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Scene-Modulated High-Order Statistical Representation Learning for No-Reference Super-Resolution Image Quality AssessmentabstractWith the rapid development of single image super-resolution (SR) technology, there is an urgent need to develop a fair no reference Super-Resolution image Quality Assessment (SRQA) method. Existing no reference SRQA methods primarily concentrate on SR artifacts including structural distortion and texture distortion by extracting spatial features, but ignore the inductive bias of Deep Neural Network (DNN)-based SR models. As a result, they function effectively for interpolation-based and dictionary-based algorithms, but struggle to perform as effectively with DNN-based SR algorithms. We found that the visual content generated by DNN-based SR models under different inductive biases often carries a content-invariant model-specific style, which can be captured by the correlations between hierarchical representation channels. To that end, we propose a novel Scene-modulated High-order Statistical Representation network (SmHSR) built on a multi-scale over-complete transformation. We quantify the perceptual quality of SR images as the shift of high-order statistical properties in their multi-scale over-complete representation, where intra-channel statistics are used to capture spatial correlations and inter-channel statistics are used to capture the inductive bias of SR models. In addition, the scene information implicit in the deep over-complete representation is used to modulate the high-order statistical properties, which simulates the top-down regulation of cognition on perception. Under the modulation of scene information, SmHSR can learn more sophisticated scene-aware statistical representation. The MultiLayer Perceptron (MLP) is used to map the high-order statistical representation to an overall quality. We test our method on multiple SR image quality databases. Experimental results show that our method outperforms the state-of-the-art SRQA methods. Yongwei Mao, Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Distilling Hierarchical Knowledge From Multimodal Fusion for Unimodal Image SegmentationabstractThe application of multimodal image fusion has become increasingly widespread across various fields in the era of deep learning. Existing fusion methods integrate infrared and visible images to provide complementary content and enhance the robustness of complex real-world scenes for high-level visual tasks, such as semantic segmentation and object detection. In return, high-level visual tasks facilitate the fusion of infrared and visible by providing mid-level semantic information. However, such frameworks rely heavily on multimodal data and require strict registration of images from different modalities before fusion, seriously limiting their practical applications due to the common realistic situations of missing modalities or misregistration. To move beyond this limitation, we propose a novel hierarchical knowledge distillation (HKD) framework tailored for unimodal image segmentation with the guidance of multi-modality. This framework aims to retain as much diverse information from multimodal image fusion as possible, thereby enhancing downstream high-level visual tasks when only the unimodal images are available during the inference phase. Our proposed method is two-stage, and we construct a robust multimodal fusion and segmentation interaction network in the first stage as a powerful teacher model. In the second stage, we design a hierarchical distillation method to transfer the fused and segmented multi-layer knowledge from the multimodal teacher model to the unimodal student model. Extensive experimental results on two public datasets, i.e., MFNet and FMB, demonstrate that the proposed hierarchical knowledge distillation framework can effectively transfuse multimodal knowledge into the unimodal student model for image enhancement and segmentation under incomplete multimodal conditions, and achieves considerably competitive results compared to multimodal image fusion and segmentation models. Weisheng Dong, Shuaibo Wang, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | CDS-Net: Contextual Difference Sensitivity Network for Pixel-Wise Road Crack DetectionabstractRoad crack detection is a key computer vision task that identifies and locates cracks in road surface images, which usually have an irregular shape and contain only a few pixels in width. Generative and unsupervised methods are popular these years, but generative methods require a lot of training data and computational power while unsupervised methods are not so satisfactory in pixel-level segmentation. The process is challenged by the irregularity of crack shapes and complex road image backgrounds. To alleviate these problems, we propose a novel method in this paper, CDS-Net, that significantly improves road crack detection performance through multiple practical modules, including the Multi-Directional Hierarchical Attention (MDHA) module and the Difference Sensitivity Reconstruction Block (DSRB). Specifically, the MDHA module employs a multi-directional feature extraction strategy to capture detailed information of cracks, thereby enhancing the discriminative power of the features. The DSRB module, designed to address the inefficiency of traditional skip-connections, utilizes masked convolution and graph convolution attention to reconstruct and refine feature representations. Additionally, we propose an improved weighted cross-entropy loss function to address the inherent class imbalance problem in road crack detection. Extensive experiments on five public datasets demonstrate that CDS-Net achieves superior performance compared to other state-of-the-art methods, showcasing its effectiveness and robustness in road crack detection. It also has a stronger generalization ability compared with other methods. Code is available athttps://github.com/ttttqz/CDS-Net/tree/master. Qinzhong Tan, Weisheng Dong, Xin Li 0005, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Discriminative Correspondence Estimation for Unsupervised RGB-D Point Cloud RegistrationabstractPoint cloud registration is a fundamental task for estimating the rigid transformation matrix between two point clouds, and is regarded as a prerequisite for downstream vision tasks. Recent works have sought to address the registration problem using the obtainable RGB-D sequence, rather than relying solely on point clouds, which may not always be available. However, most existing unsupervised RGB-D point cloud registration works struggle to obtain fine-grained, robust, discriminative correspondences due to the simple concatenation of multimodal features and the increase in vector dimensions. These methods typically follow a common paradigm: extracting features from the input data, estimating correspondences, and obtaining the transformation matrix through geometric fitting. In this work, we design a generative feature extraction module to fully leverage multimodal information, and seek a novel perspective for correspondence estimation which expands the points in the source and target point clouds into hyperrectangle-based embeddings and considers their inner relationships, based on intersections in n-dimensional space, as the basis for estimating correspondences. Each hyperrectangle-based embedding is built upon the natural and discriminative semantics from the proposed generative feature extraction module, which involves a diffusion branch, a geometric branch, and point-pixel fusion. We harness the capability of the generative model to fully leverage the information from both complementary modalities in RGB-D frames. Furthermore, this distinctive geometry space allows for efficient calculation of intersection volumes and model conditional probabilistics for estimating correspondences. Extensive experiments on the 3DMatch and ScanNet datasets show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be released at:https://github.com/cbyan1003/DCE. Chenbo Yan, Mingtao Feng, Yulan Guo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | SA-MixNet: Structure-Aware Mixup and Invariance Learning for Scribble-Supervised Road Extraction in Remote Sensing ImagesabstractMainstreamed weakly supervised road extractors rely on highly confident pseudo-labels propagated from scribbles, and their performance often degrades gradually as the image scenes tend to vary. We argue that such degradation is due to the poor model’s invariance to scenes with different complexities, whereas existing solutions to this problem are commonly based on crafted priors that cannot be derived from scribbles. To eliminate the reliance on such priors, we propose a novel structure-aware mixup and invariance learning framework (SA-MixNet) for weakly supervised road extraction that improves the model invariance in a data-driven manner. Specifically, we design a structure-aware mixup (SA-Mix) scheme to paste road regions from one image onto another to create an image scene with increased complexity while preserving the road’s structural integrity. Then, an invariance regularization is imposed on the predictions of constructed and origin images to minimize their conflicts, which thus forces the model to behave consistently in various scenes. Moreover, a discriminator-based regularization is designed to enhance connectivity while preserving the structure of roads. Combining these designs, our framework demonstrates superior performance on the DeepGlobe, Wuhan, and Massachusetts datasets, outperforming the state-of-the-art techniques by 1.47%, 2.12%, and 4.09%, respectively, in IoU metrics, and showing its potential as a plug-and-play solution. Our source code is available athttps://github.com/xdu-jjgs. Jie Feng 0003, Junpeng Zhang 0002, Weisheng Dong, Dingwen Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Locally Aware Visual State Space for Small Defect Segmentation in Complex Component ImagesabstractSegmenting small defects within large imaging fields remains challenging in industrial scenarios due to the difficulty in distinguishing defects from complex component backgrounds and identifying defects comprising only a few pixels in high-resolution images. To address these issues, we propose a novel dual-branch feature extraction architecture, the locally aware visual state space block, which captures global contextual information while maintaining locally aware perception. In addition, we introduce the parallel quad-directional scanning fusion module to extract multiscale information, aggregating high-level features at different scales for enhanced global information fusion. To avoid losing small target details when upsampling the global segmentation mask to high-resolution input size, we develop progressive location refinement modules to incrementally refine small defect localization from the bottom up. Extensive experiments on our proposed small defect segmentation dataset and a public PCB dataset demonstrate that our method outperforms existing state-of-the-art methods in both performance and efficiency. Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | CrossEI: Boosting Motion-Oriented Object Tracking With an Event CameraabstractWith the differential sensitivity and high time resolution, event cameras can record detailed motion clues, which form a complementary advantage with frame-based cameras to enhance the object tracking, especially in challenging dynamic scenes. However, how to better match heterogeneous event-image data and exploit rich complementary cues from them still remains an open issue. In this paper, we align event-image modalities by proposing a motion adaptive event sampling method, and we revisit the cross-complementarities of event-image data to design a bidirectional-enhanced fusion framework. Specifically, this sampling strategy can adapt to different dynamic scenes and integrate aligned event-image pairs. Besides, we design an image-guided motion estimation unit for extracting explicit instance-level motions, aiming at refining the uncertain event clues to distinguish primary objects and background. Then, a semantic modulation module is devised to utilize the enhanced object motion to modify the image features. Coupled with these two modules, this framework learns both the high motion sensitivity of events and the full texture of images to achieve more accurate and robust tracking. The proposed method is easily embedded in existing tracking pipelines, and trained end-to-end. We evaluate it on four large benchmarks, i.e. FE108, VisEvent, FE240hz and CoeSot. Extensive experiments demonstrate our method achieves state-of-the-art performance, and large improvements are pointed as contributions by our sampling strategy and fusion concept. Zhiwen Chen 0002, Jinjian Wu, Weisheng Dong, Leida Li, Guangming Shi |
IEEE Trans. Image Process. | 3 |
| 2025 | Cross-Frequency Attention and Color Contrast Constraint for Remote Sensing DehazingabstractCurrent deep learning-based methods for remote sensing image dehazing have developed rapidly, yet they still commonly struggle to simultaneously preserve fine texture details and restore accurate colors. The fundamental reason lies in the insufficient modeling of high-frequency information that captures structural details, as well as the lack of effective constraints for color restoration. To address the insufficient modeling of global high-frequency information, we first develop an omni-directional high-frequency feature in painting mechanism that leverages the wavelet transform to extract multi-directional high-frequency components. While maintaining the advantage of linear complexity, it models global long-range texture dependencies through cross-frequency perception. Then, to further strengthen local high-frequency representation, we design a high-frequency prompt attention module that dynamically injects wavelet-domain optimized high-frequency features as cross-level guidance signals, significantly enhancing the model's capability in edge sharpness restoration and texture detail reconstruction. Further, to alleviate the problem of inaccurate color restoration, we propose a color contrast loss function based on the HSV color space, which explicitly models the statistical distribution differences of brightness and saturation in hazy regions, guiding the model to generate dehazed images with consistent colors and natural visual appearance. Finally, extensive experiments on multiple benchmark datasets demonstrate that the proposed method outperforms existing approaches in both texture detail restoration and color consistency. Further results and code are available at: https://github.com/fyxnl/C4RSD. Jufeng Li, Yakun Ju, Chunxu Li, Weisheng Dong, Alex Chichung Kot |
IEEE Trans. Image Process. | 7 |
| 2025 | Incomplete Modalities Restoration via Hierarchical Adaptation for Robust Multimodal SegmentationabstractMultimodal semantic segmentation has significantly advanced the field of semantic segmentation by integrating data from multiple sources. However, this task often encounters missing modality scenarios due to challenges such as sensor failures or data transmission errors, which can result in substantial performance degradation. Existing approaches to addressing missing modalities predominantly involve training separate models tailored to specific missing scenarios, typically requiring considerable computational resources. In this paper, we propose a Hierarchical Adaptation framework to Restore Missing Modalities for Multimodal segmentation (HARM3), which enables frozen pretrained multimodal models to be directly applied to missing-modality semantic segmentation tasks with minimal parameter updates. Central to HARM3 is a text-instructed missing modality prompt module, which learns multimodal semantic knowledge by utilizing available modalities and textual instructions to generate prompts for the missing modalities. By incorporating a small set of trainable parameters, this module effectively facilitates knowledge transfer between high-resource domains and low-resource domains where missing modalities are more prevalent. Besides, to further enhance the model's robustness and adaptability, we introduce adaptive perturbation training and an affine modality adapter. Extensive experimental results demonstrate the effectiveness and robustness of HARM3 across a variety of missing modality scenarios. Weisheng Dong, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Image Process. | 2 |
| 2025 | Local Uncertainty Energy Transfer for Active Domain AdaptationabstractActive Domain Adaptation (ADA) improves knowledge transfer efficiency from the labeled source domain to the unlabeled target domain by selecting a few target sample labels. However, most existing active sampling methods ignore the local uncertainty of neighbors in the target domain,making it easier to pick out anomalous samples that are detrimental to the model. To address this problem, we present a new approach to active domain adaptation called Local Uncertainty Energy Transfer (LUET), which integrates active learning of local uncertainty confusion and energy transfer alignment constraints into a unified framework. First, in the active learning module, the uncertainty difficult and representative samples from the target domain are selected through local uncertainty energy selection and entropy-weighted class confusion selection. And the active learning strategy based on local uncertainty energy will avoid selecting anomalous samples in the target domain. Second, for the discrimination issue caused by domain shift, we use a global and local energy-transfer alignment constraint module to eliminate the domain gap and improve accuracy. Finally, we used negative log-likelihood loss for supervised learning of source domains and query samples. With the introduction of sample-based energy metrics, the active learning strategy is more closely with the domain alignment. Experiments on multiple domain-adaptive datasets have demonstrated that our LUET can achieve outstanding results and outperform existing state-of-the-art approaches. Guangming Shi, Weisheng Dong, Xin Li 0005, Xuemei Xie |
IEEE Trans. Image Process. | 3 |
| 2025 | Joint Spatial and Frequency Domain Learning for Lightweight Spectral Image DemosaicingabstractConventional spectral image demosaicing algorithms rely on pixels' spatial or spectral correlations for reconstruction. Due to the missing data in the multispectral filter array (MSFA), the estimation of spatial or spectral correlations is inaccurate, leading to poor reconstruction results, and these algorithms are time-consuming. Deep learning-based spectral image demosaicing methods directly learn the nonlinear mapping relationship between 2D spectral mosaic images and 3D multispectral images. However, these learning-based methods focused only on learning the mapping relationship in the spatial domain, but neglected valuable image information in the frequency domain, resulting in limited reconstruction quality. To address the above issues, this paper proposes a novel lightweight spectral image demosaicing method based on joint spatial and frequency domain information learning. First, a novel parameter-free spectral image initialization strategy based on the Fourier transform is proposed, which leads to better initialized spectral images and eases the difficulty of subsequent spectral image reconstruction. Furthermore, an efficient spatial-frequency transformer network is proposed, which jointly learns the spatial correlations and the frequency domain characteristics. Compared to existing learning-based spectral image demosaicing methods, the proposed method significantly reduces the number of model parameters and computational complexity. Extensive experiments on simulated and real-world data show that the proposed method notably outperforms existing spectral image demosaicing methods. Xun Cao, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2025 | S4DL: Shift-Sensitive Spatial-Spectral Disentangling Learning for Hyperspectral Image Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) techniques, extensively studied in hyperspectral image (HSI) classification, aim to use labeled source domain data and unlabeled target domain data to learn domain invariant features for cross-scene classification. Compared to natural images, numerous spectral bands of HSIs provide abundant semantic information, but they also increase the domain shift significantly. In most existing methods, both explicit alignment and implicit alignment simply align feature distribution, ignoring domain information in the spectrum. We noted that when the spectral channel between source and target domains is distinguished obviously, the transfer performance of these methods tends to deteriorate. Additionally, their performance fluctuates greatly owing to the varying domain shifts across various datasets. To address these problems, a novel shift-sensitive spatial-spectral disentangling learning (S4DL) approach is proposed. In S4DL, gradient-guided spatial-spectral decomposition (GSSD) is designed to separate domain-specific and domain-invariant representations by generating tailored masks under the guidance of the gradient from domain classification. A shift-sensitive adaptive monitor is defined to adjust the intensity of disentangling according to the magnitude of domain shift. Furthermore, a reversible neural network is constructed to retain domain information that lies not only in semantic but also the shallow-level detailed information. Extensive experimental results on several cross-scene HSI datasets consistently verified that S4DL is better than the state-of-the-art UDA methods. Our source code will be available athttps://github.com/xdu-jjgs/IEEE_TNNLS_S4DL. Jie Feng 0003, Junpeng Zhang 0002, Ronghua Shang, Weisheng Dong, Guangming Shi, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Pixel-Level Noise Mining for Weakly Supervised Salient Object DetectionabstractTraining a deep model for visual saliency detection requires the collection and labor-intensive annotation of overwhelmingly large data. We propose to learn saliency detection in a weakly supervised manner from single noisy label, which is easy to obtain from unsupervised handcrafted feature-based methods. However, deep networks tend to overfit such noises leading to a dramatic drop in accuracy. Given our goal, we address a natural question: can we identify outliers during network prediction and rectify the label noises? To this end, we propose a pixel-level noise mining framework for robust salient object detection (SOD) by exploiting its own knowledge, and without the need for external models. Specifically, during the early training stage, we progressively identify the outliers from a novel perspective during saliency detection, before the network overfits to the noisy labels, and generate a selection matrix in each iteration. Next, we adaptively rectify the label noises under the guidance of the selection matrix for better supervision in the later training stage. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our method showing its ability to learn saliency detection comparable to state-of-the-art fully supervised methods. Furthermore, our approach outperforms existing weakly supervised methods utilizing single noisy label and surpasses the half of existing weakly supervised methods employing multiple noisy labels. Our approach, which trains with multiple noisy labels, outperforms all other methods employing multiple noisy labels across four major datasets. Furthermore, we also evaluate the generalization ability of our method on the multiclass semantic segmentation (SS) task. Our code is available at https://github.com/kendongdong/NoiseMining. Kendong Liu, Mingtao Feng, Wei Zhao 0019, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Inverse Weight-Balancing for Deep Long-Tailed LearningabstractThe performance of deep learning models often degrades rapidly when faced with imbalanced data characterized by a long-tailed distribution. Researchers have found that the fully connected layer trained by cross-entropy loss has large weight-norms for classes with many samples, but not for classes with few samples. How to address the data imbalance problem with both the encoder and the classifier seems an under-researched problem. In this paper, we propose an inverse weight-balancing (IWB) approach to guide model training and alleviate the data imbalance problem in two stages. In the first stage, an encoder and classifier (the fully connected layer) are trained using conventional cross-entropy loss. In the second stage, with a fixed encoder, the classifier is finetuned through an adaptive distribution for IWB in the decision space. Unlike existing inverse image frequency that implements a multiplicative margin adjustment transformation in the classification layer, our approach can be interpreted as an adaptive distribution alignment strategy using not only the class-wise number distribution but also the sample-wise difficulty distribution in both encoder and classifier. Experiments show that our method can greatly improve performance on imbalanced datasets such as CIFAR100-LT with different imbalance factors, ImageNet-LT, and iNaturelists2018. Wenqi Dang, Weisheng Dong, Xin Li 0005, Guangming Shi |
AAAI | 3 |
| 2024 | External Knowledge Enhanced 3D Scene Generation from Sketch
Mingtao Feng, Yaonan Wang 0001, He Xie, Weisheng Dong, Bo Miao, Ajmal Mian |
ECCV (6) | 5 |
| 2024 | Hyperspectral Image Denoising Based on Nonlocal Memory-Augmented Spectral AttentionabstractIn this paper, we propose an effective denoising method aimed at restoring clean images from hyperspectral images disturbed by multiple types of noise. Previously, data-based hyperspectral denoising methods mainly utilize adjacent spectral bands to assist in image restoration. However, due to the limited information provided by adjacent bands under the same level of noise interference, the performance of hyperspectral denoising algorithms is limited. In contrast, this paper introduces a cross-band non-local attention strategy that can supplement the spectral non-local similarity relationship of hyperspectral data. In addition, to alleviate the problem of different levels of noise interference in different bands, this paper uses memory-augmented attention to introduce global spectral information into each band, thereby fully utilizing inter band information to remove mixed noise. Furthermore, this paper uses multi-scale spatial information combined with spectral information to enhance denoising performance. The experiments conducted on both synthetic and real data have demonstrated the superiority of the proposed method over other methods. Yige Mo, Weisheng Dong |
IGARSS | 5 |
| 2024 | Semantics-Aware Image Aesthetics Assessment using Tag Matching and Contrastive RankingabstractThe perception of image aesthetics is built upon the understanding of semantic content. However, how to evaluate the aesthetic quality of images with diversified semantic backgrounds remains challenging in image aesthetics assessment (IAA). To address the dilemma, this paper presents a semantics-aware image aesthetics assessment approach, which first analyzes the semantic content of images and then models the aesthetic distinctions among images from two perspectives, i.e., aesthetic attribute and aesthetic level. Concretely, we propose two strategies, dubbed tag matching and contrastive ranking, to extract knowledge pertaining to image aesthetics. The tag matching identifies the semantic category and the dominant aesthetic attributes based on predefined tag libraries. The contrastive ranking is designed to uncover the comparative relationships among images with different aesthetic levels but similar semantic backgrounds. In the process of contrastive ranking, the impact of long-tailed distribution of aesthetic data is also considered by balanced sampling and traversal contrastive learning. Extensive experiments and comparisons on three benchmark IAA databases demonstrate the superior performance of the proposed model in terms of both prediction accuracy and alleviating long-tailed effect. The code will be public at https://github.com/yzc-ippl/TMCR **REMOVE 2nd URL**://github.com/yzc-ippl/TMCR. Zhichao Yang 0013, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong |
ACM Multimedia | 5 |
| 2024 | TSUDepth: Exploring temporal symmetry-based uncertainty for unsupervised monocular depth estimation
Yufan Zhu, Weisheng Dong, Xin Li 0005, Guangming Shi |
Neurocomputing | 3 |
| 2024 | Quality-aware blind image motion deblurring
Tianshu Song, Leida Li, Jinjian Wu, Weisheng Dong, Deqiang Cheng 0001 |
Pattern Recognit. | 4 |
| 2024 | Learning real-world heterogeneous noise models with a benchmark dataset
Jie Lin 0008, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
Pattern Recognit. | 3 |
| 2024 | 3D Object Detection From Point Cloud via Voting Step Diffusionabstract3D object detection is a fundamental task in scene understanding. Numerous research efforts have been dedicated to better incorporate Hough voting into the 3D object detection pipeline. However, due to the noisy, cluttered, and partial nature of real 3D scans, existing voting-based methods often receive votes from the partial surfaces of individual objects together with severe noises, leading to sub-optimal detection performance. In this work, we focus on the distributional properties of point clouds and formulate the voting process as generating new points in the high-density region of the distribution of object centers. To achieve this, we propose a new method to move random 3D points toward the high-density region of the distribution by estimating the score function of the distribution with a noise conditioned score network. Specifically, we first generate a set of object center proposals to coarsely identify the high-density region of the object center distribution. To estimate the score function, we perturb the generated object center proposals by adding normalized Gaussian noise, and then jointly estimate the score function of all perturbed distributions. Finally, we generate new votes by moving random 3D points to the high-density region of the object center distribution according to the estimated score function. Extensive experiments on two large scale indoor 3D scene datasets, SUN RGB-D and ScanNet V2, demonstrate the superiority of our proposed method. The code will be released athttps://github.com/HHrEtvP/DiffVote. Haoran Hou, Mingtao Feng, Weisheng Dong, Qing Zhu 0003, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Motion-Oriented Hybrid Spiking Neural Networks for Event-Based Motion DeblurringabstractImage deblurring based only on the blurry image is challenging as motion information is lost while imaging. Event cameras capture the texture of moving objects in high temporal resolution with asynchronous events. In this paper, we extract motion features from events and fuse them with background features from the image for event-based image deblurring. Spiking neural network (SNN), a widely recognized event feature extractor, is well suited for motion feature extraction due to its high temporal resolution. However, extracting motion information from events exclusively with SNN is challenging. We propose a novel Temporal-local-Spatio Spiking Transformer (TSST) to extract motion intensity and motion attention regions in the spatio-temporal domain. Motion intensity extracted from spiking features is represented as a high temporal resolution motion attention map to guide the fusion of the two networks. In the temporal domain, motion intensity maps spiking features to CNN features as motion features to avoid blurring. In the spatial domain, the motion intensity shows the motion regions and gives the weight of the motion feature during fusion. Moreover, a hybrid feature extraction encoder (HFEE) is introduced, which fully fuses the motion and background features for deblurring. The gradient is back-propagated from CNN to SNN, and the hybrid deblurring network is jointly optimized. We evaluated the performance of our model on the public dataset GoPro and a real event dataset we captured. Codes and pretrained models are available athttps://github.com/XDULzx/MotionSNN. Zhaoxin Liu, Jinjian Wu, Guangming Shi, Wen Yang 0008, Weisheng Dong, Qinghang Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Uncertainty Modeling of the Transmission Map for Single Image DehazingabstractDespite rapid progress of end-to-end optimization for single-image dehazing, a long-standing open problem is the non-homogenous haze, at the core of the differences between synthetic hazy images and real hazy images. The atmospheric scattering model (ASM) has been widely adopted to model the degradation process of haze images but based on the assumption of homogeneous haze. In realistic scenarios, non-homogeneous haze often makes it more difficult to estimate the transmission map in ASM, resulting in undesired artifacts in the restored images. To address the issue of non-homogeneous haze, we propose to model the uncertainty in the estimation of the transmission map and develop a spatially adaptive learning module for ASM correction. Specifically, we present an approach to enhancing the well-known Dark Channel prior (DCP) by relaxing the constraint with the transmission map in the DCP-net. Assuming the availability of paired training data, we have developed a strategy to address vulnerability in the DCP, leading to a more accurate estimation of the transmission map. Then, we explore the uncertainty between the estimated transmission map and target transmission map (Ground Truth) to reformulate the ASM for the presence of non-homogeneous haze. A robust and accurate estimated transmission map can boost the final dehazing performance of our DCP-net. Experiments on three popular synthetic and real non-homogeneous datasets show that our proposed approach has achieved better results on both synthetic scenes and real non-homogeneous scenes. The code is available athttps://see.xidian.edu.cn/faculty/wsdong/Projects/Projects/project_dehazing_TCSVT2024.htm Bokang Wang, Qian Ning, Xin Li 0005, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | TransVQA: Transferable Vector Quantization Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the unlabeled target domain. Most existing domain adaptation methods are based on convolutional neural networks (CNNs) to learn cross-domain invariant features. Inspired by the success of transformer architectures and their superiority to CNNs, we propose to combine the transformer with UDA to improve their generalization properties. In this paper, we present a novel model named Trans ferable V ector Q uantization A lignment for Unsupervised Domain Adaptation (TransVQA), which integrates the Transferable transformer-based feature extractor (Trans), vector quantization domain alignment (VQA), and mutual information weighted maximization confusion matrix (MIMC) of intra-class discrimination into a unified domain adaptation framework. First, TransVQA uses the transformer to extract more accurate features in different domains for classification. Second, TransVQA, based on the vector quantization alignment module, uses a two-step alignment method to align the extracted cross-domain features and solve the domain shift problem. The two-step alignment includes global alignment via vector quantization and intra-class local alignment via pseudo-labels. Third, for intra-class feature discrimination problem caused by the fuzzy alignment of different domains, we use the MIMC module to constrain the target domain output and increase the accuracy of pseudo-labels. The experiments on several datasets of domain adaptation show that TransVQA can achieve excellent performance and outperform existing state-of-the-art methods. Weisheng Dong, Xin Li 0005, Guangming Shi, Xuemei Xie |
IEEE Trans. Image Process. | 2 |
| 2024 | Multi-Scale Spatio-Temporal Memory Network for Lightweight Video DenoisingabstractDeep learning-based video denoising methods have achieved great performance improvements in recent years. However, the expensive computational cost arising from sophisticated network design has severely limited their applications in real-world scenarios. To address this practical weakness, we propose a multiscale spatio-temporal memory network for fast video denoising, named MSTMN, aiming at striking an improved trade-off between cost and performance. To develop an efficient and effective algorithm for video denoising, we exploit a multiscale representation based on the Gaussian-Laplacian pyramid decomposition so that the reference frame can be restored in a coarse-to-fine manner. Guided by a model-based optimization approach, we design an effective variance estimation module, an alignment error estimation module and an adaptive fusion module for each scale of the pyramid representation. For the fusion module, we employ a reconstruction recurrence strategy to incorporate local temporal information. Moreover, we propose a memory enhancement module to exploit the global spatio-temporal information. Meanwhile, the similarity computation of the spatio-temporal memory network enables the proposed network to adaptively search the valuable information at the patch level, which avoids computationally expensive motion estimation and compensation operations. Experimental results on real-world raw video datasets have demonstrated that the proposed lightweight network outperforms current state-of-the-art fast video denoising algorithms such as FastDVDnet, EMVD, and ReMoNet with fewer computational costs. Xin Li 0005, Jie Lin 0008, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 6 |
| 2024 | Learning Frame-Event Fusion for Motion DeblurringabstractMotion deblurring is a highly ill-posed problem due to the significant loss of motion information in the blurring process. Complementary informative features from auxiliary sensors such as event cameras can be explored for guiding motion deblurring. The event camera can capture rich motion information asynchronously with microsecond accuracy. In this paper, a novel frame-event fusion framework is proposed for event-driven motion deblurring (FEF-Deblur), which can sufficiently explore long-range cross-modal information interactions. Firstly, different modalities are usually complementary and also redundant. Cross-modal fusion is modeled as complementary-unique features separation-and-aggregation, avoiding the modality redundancy. Unique features and complementary features are first inferred with parallel intra-modal self-attention and inter-modal cross-attention respectively. After that, a correlation-based constraint is designed to act between unique and complementary features to facilitate their differentiation, which assists in cross-modal redundancy suppression. Additionally, spatio-temporal dependencies among neighboring inputs are crucial for motion deblurring. A recurrent cross attention is introduced to preserve inter-input attention information, in which the current spatial features and aggregated temporal features are attending to each other by establishing the long-range interaction between them. Extensive experiments on both synthetic and real-world motion deblurring datasets demonstrate our method outperforms state-of-the-art event-based and image/video-based methods. The code will be made publicly available. Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2024 | Blind Image Quality Assessment Based on Perceptual ComparisonabstractBlind image quality assessment (BIQA) is a regression task with continuous label space, the feature space of which is expected to have a corresponding continuity in the target space. However, existing approaches typically learn quality score regression directly in an end-to-end fashion, which leaves networks susceptible to interference from task-agnostic information, and fails to capture the continuity of BIQA. In this work, by explicitly establishing inter-sample associations, a simple yet effective BIQA framework based on perceptual comparison is proposed to capture the continuity. To this end, besides the basic quality score regression, the relative quality scores between images are predicted to exploit the relative quality relationships between samples for optimizing the representation of image perceptual quality. In addition, based on the human perceptual characteristic, we derive a novel sample weighting strategy to dynamically adjust the weights for different samples in the network learning process for further improving the robustness of the model. The performances on both single-database and cross-database experiments achieve state-of-the-art, indicating the effectiveness of the proposed method. Besides, the proposed framework is model-agnostic, which can effectively improve the performance of the benchmark model with no extra inference cost. Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Multim. | 5 |
| 2023 | Self-supervised Non-uniform Kernel Estimation with Flow-based Motion Prior for Blind Image DeblurringabstractMany deep learning-based solutions to blind image deblurring estimate the blur representation and reconstruct the target image from its blurry observation. However, these methods suffer from severe performance degradation in real-world scenarios because they ignore important prior information about motion blur (e.g., real-world motion blur is diverse and spatially varying). Some methods have attempted to explicitly estimate non-uniform blur kernels by CNNs, but accurate estimation is still challenging due to the lack of ground truth about spatially varying blur kernels in real-world images. To address these issues, we propose to represent the field of motion blur kernels in a latent space by normalizing flows, and design CNNs to predict the latent codes instead of motion kernels. To further improve the accuracy and robustness of non-uniform kernel estimation, we introduce uncertainty learning into the process of estimating latent codes and propose a multi-scale kernel attention module to better integrate image features with estimated kernels. Extensive experimental results, especially on real-world blur datasets, demonstrate that our method achieves state-of-the-art results in terms of both subjective and objective quality as well as excellent generalization performance for non-uniform image deblurring. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UFPNet.htm. Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
CVPR | 3 |
| 2023 | Vector Quantization with Self-Attention for Quality-Independent Representation LearningabstractRecently, the robustness of deep neural networks has drawn extensive attention due to the potential distribution shift between training and testing data (e.g., deep models trained on high-quality images are sensitive to corruption during testing). Many researchers attempt to make the model learn invariant representations from multiple corrupted data through data augmentation or image-pair-based feature distillation to improve the robustness. Inspired by sparse representation in image restoration, we opt to address this issue by learning image-quality-independent feature representation in a simple plug-and-play manner, that is, to introduce discrete vector quantization (VQ) to remove redundancy in recognition models. Specifically, we first add a codebook module to the network to quantize deep features. Then we concatenate them and design a self-attention module to enhance the representation. During training, we enforce the quantization of features from clean and corrupted images in the same discrete embedding space so that an invariant quality-independentfeature representation can be learned to improve the recognition robustness of low-quality images. Qualitative and quantitative experimental results show that our method achieved this goal effectively, leading to a new state-of-the-art result of 43.1 % mCE on ImageNet-C with ResNet50 as the backbone. On other robustness benchmark datasets, such as ImageNet-R, our method also has an accuracy improvement of almost 2%. The source code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/VQSA.htm Weisheng Dong, Xin Li 0005, Mengluan Huang, Guangming Shi |
CVPR | 2 |
| 2023 | Low-Light Image Enhancement with Multi-stage Residue Quantization and Brightness-aware AttentionabstractLow-light image enhancement (LLIE) aims to recover illumination and improve the visibility of low-light images. Conventional LLIE methods often produce poor results because they neglect the effect of noise interference. Deep learning-based LLIE methods focus on learning a mapping function between low-light images and normal-light images that outperforms conventional LLIE methods. However, most deep learning-based LLIE methods cannot yet fully exploit the guidance of auxiliary priors provided by normal-light images in the training dataset. In this paper, we propose a brightness-aware network with normal-light priors based on brightness-aware attention and residual-quantized codebook. To achieve a more natural and realistic enhancement, we design a query module to obtain more reliable normal-light features and fuse them with low-light features by a fusion branch. In addition, we propose a brightness-aware attention module to further improve the robustness of the network to the brightness. Extensive experimental results on both real-captured and synthetic data show that our method outperforms existing state-of-the-art methods. Weisheng Dong, Xin Li 0005, Guangming Shi |
ICCV | 3 |
| 2023 | Exploring Correlations in Degraded Spatial Identity Features for Blind Face RestorationabstractBlind face restoration aims to recover high-quality face images from low-quality ones with complex and unknown degradation. Existing approaches have achieved promising performance by leveraging pre-trained dictionaries or generative priors. However, these methods may fail to exploit the full potential of degraded inputs and facial identity features due to complex degradation. To address this issue, we propose a novel method that explores the correlation of degraded spatial identity features by learning a general representation using memory network. Specifically, our approach enhances degraded features with more identity by leveraging similar facial features retrieved from memory network. We also propose a fusion approach that fuses memorized spatial features with GAN prior features via affine transformation and blending fusion to improve fidelity and realism. Additionally, the memory network is updated online in an unsupervised manner along with other modules, which obviates the requirement for pre-training. Experimental results on synthetic and popular real-world datasets demonstrate the effectiveness of our proposed method, which achieves at least comparable and often better performance than other state-of-the-art approaches. Qian Ning, Weisheng Dong, Xin Li 0005, Guangming Shi |
ACM Multimedia | 3 |
| 2023 | AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentabstractImage aesthetics assessment (IAA) aims at predicting the aesthetic quality of images. Recently, large pre-trained vision-language models, like CLIP, have shown impressive performances on various visual tasks. When it comes to IAA, a straightforward way is to finetune the CLIP image encoder using aesthetic images. However, this can only achieve limited success without considering the uniqueness of multimodal data in the aesthetics domain. People usually assess image aesthetics according to fine-grained visual attributes, e.g., color, light and composition. However, how to learn aesthetics-aware attributes from CLIP-based semantic space has not been addressed before. With this motivation, this paper presents a CLIP-based multi-attribute contrastive learning framework for IAA, dubbed AesCLIP. Specifically, AesCLIP consists of two major components, i.e., aesthetic attribute-based comment classification and attribute-aware learning. The former classifies the aesthetic comments into different attribute categories. Then the latter learns an aesthetic attribute-aware representation by contrastive learning, aiming to mitigate the domain shift from the general visual domain to the aesthetics domain. Extensive experiments have been done by using the pre-trained AesCLIP on four popular IAA databases, and the results demonstrate the advantage of AesCLIP over the state-of-the-arts. The source code will be public at https://github.com/OPPOMKLab/AesCLIP. Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong, Yuzhe Yang 0001, Liwu Xu, Guangming Shi |
ACM Multimedia | 5 |
| 2023 | Event-based Motion Deblurring with Modality-Aware Decomposition and RecompositionabstractEvent camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events. Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 4 |
| 2023 | Memory Based Temporal Fusion Network for Video Deblurring
Chaohua Wang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
Int. J. Comput. Vis. | 2 |
| 2023 | Transfer learning for just noticeable difference estimation
Yongwei Mao, Jinjian Wu, Leida Li, Weisheng Dong |
Inf. Sci. | 5 |
| 2023 | MADPL-net: Multi-layer attention dictionary pair learning network for image classification
Guangming Shi, Weisheng Dong, Xuemei Xie |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Deep Gaussian Scale Mixture Prior for Image ReconstructionabstractImage reconstruction from partial observations has attracted increasing attention. Conventional image reconstruction methods with hand-crafted priors often fail to recover fine image details due to the poor representation capability of the hand-crafted priors. Deep learning methods attack this problem by directly learning mapping functions between the observations and the targeted images can achieve much better results. However, most powerful deep networks lack transparency and are nontrivial to design heuristically. This paper proposes a novel image reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Unlike existing unfolding methods that only estimate the image means (i.e., the denoising prior) but neglected the variances, we propose characterizing images by the GSM models with learned means and variances through a deep network. Furthermore, to learn the long-range dependencies of images, we develop an enhanced variant based on the Swin Transformer for learning GSM models. All parameters of the MAP estimator and the deep network are jointly optimized through end-to-end training. Extensive simulation and real data experimental results on spectral compressive imaging and image super-resolution demonstrate that the proposed method outperforms existing state-of-the-art methods. Xin Yuan 0002, Weisheng Dong, Jinjian Wu, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Adaptive Search-and-Training for Robust and Efficient Network PruningabstractBoth network pruning and neural architecture search (NAS) can be interpreted as techniques to automate the design and optimization of artificial neural networks. In this paper, we challenge the conventional wisdom of training before pruning by proposing a joint search-and-training approach to learn a compact network directly from scratch. Using pruning as a search strategy, we advocate three new insights for network engineering: 1) to formulate adaptive search as a cold start strategy to find a compact subnetwork on the coarse scale; and 2) to automatically learn the threshold for network pruning; 3) to offer flexibility to choose between efficiency and robustness. More specifically, we propose an adaptive search algorithm in the cold start by exploiting the randomness and flexibility of filter pruning. The weights associated with the network filters will be updated by ThreshNet, a flexible coarse-to-fine pruning method inspired by reinforcement learning. In addition, we introduce a robust pruning strategy leveraging the technique of knowledge distillation through a teacher-student network. Extensive experiments on ResNet and VGGNet have shown that our proposed method can achieve a better balance in terms of efficiency and accuracy and notable advantages over current state-of-the-art pruning methods in several popular datasets, including CIFAR10, CIFAR100, and ImageNet. The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/AST-NP.htm. Xiaotong Lu, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Uncertainty-Driven Knowledge Distillation for Language Model CompressionabstractDespite the remarkable performance on various Natural Language Processing (NLP) tasks, the parametric complexity of pretrained language models has remained a major obstacle due to limited computational resources in many practical applications. Techniques such as knowledge distillation, network pruning, and quantization have been developed for language model compression. However, it has remained challenging to achieve an optimal tradeoff between model size and inference accuracy. To address this issue, we propose a novel and efficient uncertainty-driven knowledge distillation compression method for transformer-based pretrained language models. Specifically, we design a method of parameter retention and feedforward network parameter distillation to compress N-stacked transformer modules into one module in the fine-tuning stage. A key innovation of our approach is to add the uncertainty estimation module (UEM) into the student network such that it can guide the student network's feature reconstruction in the latent space (similar to the teacher's). Across multiple datasets in the natural language inference tasks of GLUE, we have achieved more than 95% accuracy of the original BERT, while only using about 50% of the parameters. Weisheng Dong, Xin Li 0005, Guangming Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Dynamic Expert-Knowledge Ensemble for Generalizable Video Quality AssessmentabstractDespite the impressive progress of supervised methods in quality assessment for in- the-wild videos, models trained on one domain often fail to generalize well to others due to the domain shifts caused by distortion diversity and content variation. Domain generalizable video quality assessment (VQA) methods that can work across domains remain an open research challenge. Although combining more data following the mixed-domain training strategy can improve the generalization performance to a certain extent, the specific knowledge from each source domain, which could potentially be useful for improving unseen domain generalization, is ignored in this principle. Motivated by this, we propose a domain generalizable VQA method named Dynamic Ensemble of Expert-Knowledge (DEEK), a novel framework that dynamically exploits the expert-knowledge from each source domain to achieve a generalizable ensemble prediction. Specifically, based on the multiple experts each trained to specialize in a particular source domain, we aim to exploit complementary information provided by the expert-knowledge. We effectively train an ensemble model by proposing a quality-sensitive InfoNCE loss to regularize the collaborative training of all experts in the contrastive learning formulation, aiming to exploit complementary information provided by the expert-knowledge when forming the ensemble. By dynamically integrating the experts according to their relevances to the target data, these expert-knowledge could be leveraged for better generalization. Experiments on five VQA datasets verify that our approach outperforms the state-of-the-arts by large margins. Pengfei Chen 0003, Leida Li, Haoliang Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | ECSNet: Spatio-Temporal Feature Learning for Event CameraabstractThe neuromorphic event cameras can efficiently sense the latent geometric structures and motion clues of a scene by generating asynchronous and sparse event signals. Due to the irregular layout of the event signals, how to leverage their plentiful spatio-temporal information for recognition tasks remains a significant challenge. Existing methods tend to treat events as dense image-like or point-serie representations. However, they either suffer from severe destruction on the sparsity of event data or fail to encode robust spatial cues. To fully exploit their inherent sparsity with reconciling the spatio-temporal information, we introduce a compact event representation, namely 2D-1T event cloud sequence (2D-1T ECS). We couple this representation with a novel light-weight spatio-temporal learning framework (ECSNet) that accommodates both object classification and action recognition tasks. The core of our framework is a hierarchical spatial relation module. Equipped with specially designed surface-event-based sampling unit and local event normalization unit to enhance the inter-event relation encoding, this module learns robust geometric features from the 2D event clouds. And we propose a motion attention module for efficiently capturing long-term temporal context evolving with the 1T cloud sequence. Empirically, the experiments show that our framework achieves par or even better state-of-the-art performance. Importantly, our approach cooperates well with the sparsity of event data without any sophisticated operations, hence leading to low computational costs and prominent inference speeds. Zhiwen Chen 0002, Jinjian Wu, Junhui Hou, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Differentiable Neural Architecture Search for Extremely Lightweight Image Super-ResolutionabstractSingle Image Super-Resolution (SISR) tasks have achieved significant performance with deep neural networks. However, the large number of parameters in CNN-based methods for SISR tasks require heavy computations. Although several efficient SISR models have been recently proposed, most are handcrafted and thus lack flexibility. In this work, we propose a novel differentiable Neural Architecture Search (NAS) approach on both the cell-level and network-level to search for lightweight SISR models. Specifically, the cell-level search space is designed based on an information distillation mechanism, focusing on the combinations of lightweight operations and aiming to build a more lightweight and accurate SR structure. The network-level search space is designed to consider the feature connections among the cells and aims to find which information flow benefits the cell most to boost the performance. Unlike the existing Reinforcement Learning (RL) or Evolutionary Algorithm (EA) based NAS methods for SISR tasks, our search pipeline is fully differentiable, and the lightweight SISR models can be efficiently searched on both the cell-level and network-level jointly on a single GPU. Experiments show that our methods can achieve state-of-the-art performance on the benchmark datasets in terms of PSNR, SSIM, and model complexity with merely 68G Multi-Adds for$\times 2$and 18G Multi-Adds for$\times 4$SR tasks. Li Shen 0008, Chaoyang He 0001, Weisheng Dong, Wei Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Quality Assessment of UGC Videos Based on Decomposition and RecompositionabstractThe prevalence of short-video applications imposes more requirements for video quality assessment (VQA). User-generated content (UGC) videos are captured under an unprofessional environment, thus suffering from various dynamic degradations, such as camera shaking. To cover the dynamic degradations, existing recurrent neural network-based UGC-VQA methods can only provide implicit modeling, which is unclear and difficult to analyze. In this work, we consider explicit motion representation for dynamic degradations, and propose a motion-enhanced UGC-VQA method based on decomposition and recomposition. In the decomposition stage, a dual-stream decomposition module is built, and VQA task is decomposed into single frame-based quality assessment problem and cross frames-based motion understanding. The dual streams are well grounded on the two-pathway visual system during perception, and require no extra UGC data due to knowledge transfer. Hierarchical features from shallow to deep layers are gathered to narrow the gaps from tasks and domains. In the recomposition stage, a progressively residual aggregation module is built to recompose features from the dual streams. Representations with different layers and pathways are interacted and aggregated in a progressive and residual manner, which keeps a good trade-off between representation deficiency and redundancy. Extensive experiments on UGC-VQA databases verify that our method achieves the state-of-the-art performance and keeps a good capability of generalization. The source code will be available inhttps://github.com/Sissuire/DSD-PRO. Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Deep Unfolding Network for Efficient Mixed Video Noise RemovalabstractExisting image and video denoising algorithms have focused on removing homogeneous Gaussian noise. However, this assumption with noise modeling is often too simplistic for the characteristics of real-world noise. Moreover, the design of network architectures in most deep learning-based video denoising methods is heuristic, ignoring valuable domain knowledge. In this paper, we propose a model-guided deep unfolding network for the more challenging and realistic mixed noise video denoising problem, named DU-MVDnet. First, we develop a novel observation model/likelihood function based on the correlations among adjacent degraded frames. In the framework of Bayesian deep learning, we introduce a deep image denoiser prior and obtain an iterative optimization algorithm based on the maximum a posterior (MAP) estimation. To facilitate end-to-end optimization, the iterative algorithm is transformed into a deep convolutional neural network (DCNN)-based implementation. Furthermore, recognizing the limitations of traditional motion estimation and compensation methods, we propose an efficient multistage recursive fusion strategy to exploit temporal dependencies. Specifically, we divide video frames into several overlapping groups and progressively integrate these frames into one frame. Toward this objective, we implement a multiframe adaptive aggregation operation to integrate feature maps of intragroup with those of intergroup frames. Extensive experimental results on different video test datasets have demonstrated that the proposed model-guided deep network outperforms current state-of-the-art video denoising algorithms such as FastDVDnet and MAP-VDNet. Xin Li 0005, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Supervised Contrastive Learning Based on Fusion of Global and Local Features for Remote Sensing Image RetrievalabstractWith the rapid development of remote sensing sensor technology, the number of remote sensing images (RSIs) has exploded. How to effectively retrieve and manage these massive data has become an urgent problem. At present, content-based image retrieval (CBIR) methods have become a mainstream method due to their excellent performance. However, most of the existing retrieval methods only consider the global features of images, which lacks the ability to discriminate images with the same semantic information but different visual representations. To alleviate this issue, supervised contrastive learning based on the fusion of global and local features method is proposed in this paper, named SCFR. Firstly, a fusion module is designed to combine global and local features to enhance the ability of image expression. Secondly, supervised contrastive learning is introduced into the retrieval task to effectively improve the feature distribution, so that the positive sample pairs are close to each other, and the negative sample pairs are far away from each other in the feature space. Furthermore, to make the distribution of features of the same class more compact, the center contrastive loss is added to the constraints, and combines the class centers that change iteratively with the network. Experimental results on three RSI datasets show that our proposed method has more effective retrieval performance than the state-of-the-art methods. The code and models are available at https://github.com/xdplay17/SCFR. Mengluan Huang, Weisheng Dong, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Searching Efficient Model-Guided Deep Network for Image DenoisingabstractUnlike the success of neural architecture search (NAS) in high-level vision tasks, it remains challenging to find computationally efficient and memory-efficient solutions to low-level vision problems such as image restoration through NAS. One of the fundamental barriers to differential NAS-based image restoration is the optimization gap between the super-network and the sub-architectures, causing instability during the searching process. In this paper, we present a novel approach to fill this gap in image denoising application by connecting model-guided design (MoD) with NAS (MoD-NAS). Specifically, we propose to construct a new search space under a model-guided framework and develop more stable and efficient differential search strategies. MoD-NAS employs a highly reusable width search strategy and a densely connected search block to automatically select the operations of each layer as well as network width and depth via gradient descent. During the search process, the proposed MoD-NAS remains stable because of the smoother search space designed under the model-guided framework. Experimental results on several popular datasets show that our MoD-NAS method has achieved at least comparable even better PSNR performance than current state-of-the-art methods with fewer parameters, fewer flops, and less testing time. "The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/Mod-NAS.htm". Qian Ning, Weisheng Dong, Xin Li 0005, Jinjian Wu |
IEEE Trans. Image Process. | 2 |
| 2023 | Spatially Varying Prior Learning for Blind Hyperspectral Image FusionabstractIn recent years, researchers have become more interested in hyperspectral image fusion (HIF) as a potential alternative to expensive high-resolution hyperspectral imaging systems, which aims to recover a high-resolution hyperspectral image (HR-HSI) from two images obtained from low-resolution hyperspectral (LR-HSI) and high-spatial-resolution multispectral (HR-MSI). It is generally assumed that degeneration in both the spatial and spectral domains is known in traditional model-based methods or that there existed paired HR-LR training data in deep learning-based methods. However, such an assumption is often invalid in practice. Furthermore, most existing works, either introducing hand-crafted priors or treating HIF as a black-box problem, cannot take full advantage of the physical model. To address those issues, we propose a deep blind HIF method by unfolding model-based maximum a posterior (MAP) estimation into a network implementation in this paper. Our method works with a Laplace distribution (LD) prior that does not need paired training data. Moreover, we have developed an observation module to directly learn degeneration in the spatial domain from LR-HSI data, addressing the challenge of spatially-varying degradation. We also propose to learn the uncertainty (mean and variance) of LD models using a novel Swin-Transformer-based denoiser and to estimate the variance of degraded images from residual errors (rather than treating them as global scalars). All parameters of the MAP estimation algorithm and the observation module can be jointly optimized through end-to-end training. Extensive experiments on both synthetic and real datasets show that the proposed method outperforms existing competing methods in terms of both objective evaluation indexes and visual qualities. Xin Li 0005, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2022 | Robust Depth Completion with Uncertainty-Driven Loss FunctionsabstractRecovering a dense depth image from sparse LiDAR scans is a challenging task. Despite the popularity of color-guided methods for sparse-to-dense depth completion, they treated pixels equally during optimization, ignoring the uneven distribution characteristics in the sparse depth map and the accumulated outliers in the synthesized ground truth. In this work, we introduce uncertainty-driven loss functions to improve the robustness of depth completion and handle the uncertainty in depth completion. Specifically, we propose an explicit uncertainty formulation for robust depth completion with Jeffrey's prior. A parametric uncertain-driven loss is introduced and translated to new loss functions that are robust to noisy or missing data. Meanwhile, we propose a multiscale joint prediction model that can simultaneously predict depth and uncertainty maps. The estimated uncertainty map is also used to perform adaptive prediction on the pixels with high uncertainty, leading to a residual map for refining the completion results. Our method has been tested on KITTI Depth Completion Benchmark and achieved the state-of-the-art robustness performance in terms of MAE, IMAE, and IRMSE metrics. Yufan Zhu, Weisheng Dong, Leida Li, Jinjian Wu, Xin Li 0005, Guangming Shi |
AAAI | 2 |
| 2022 | Uncertainty Learning in Kernel Estimation for Multi-stage Blind Image Super-Resolution
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
ECCV (18) | 2 |
| 2022 | Self-feature Distillation with Uncertainty Modeling for Degraded Image Recognition
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
ECCV (24) | 2 |
| 2022 | Supervised Contrastive Learning-Based Deep Hash Retrieval for Remote Sensing ImageabstractWith the development of remote sensing technology, the earth observation data shows a blowout growth. How to quickly and accurately retrieve the target content from the massive data has become a noteworthy task. Recently, the methods based on convolutional neural networks (CNN) have been far ahead in remote sensing retrieval. However, due to the diversity of sources acquired, remote sensing images with the same semantic information may have great visual differences. The pre-trained CNN is not always able to successfully extract representative and distinguishing features to recognize the target content. To address this problem, a supervised contrastive learning-based deep hash retrieval method (SCLDHR) is introduced in this paper, which effectively uses label information to gather image features belonging to the same class and separate image features of different classes in the embedding space. Furthermore, a quantization function in the hash space is designed to improve hashing quality. The experimental results conducted on three datasets show that SCLDHR can achieve competitive retrieval performance compared with state-of-the-art methods. Mengluan Huang, Weisheng Dong, Guangming Shi |
IGARSS | 3 |
| 2022 | Learning Degradation Uncertainty for Unsupervised Real-world Image Super-resolutionabstractAcquiring degraded images with paired high-resolution (HR) images is often challenging, impeding the advance of image super-resolution in real-world applications. By generating realistic low-resolution (LR) images with degradation similar to that in real-world scenarios, simulated paired LR-HR data can be constructed for supervised training. However, most of the existing work ignores the degradation uncertainty of the generated realistic LR images, since only one LR image has been generated given an HR image. To address this weakness, we propose learning the degradation uncertainty of generated LR images and sampling multiple LR images from the learned LR image (mean) and degradation uncertainty (variance) and construct LR-HR pairs to train the super-resolution (SR) networks. Specifically, uncertainty can be learned by minimizing the proposed loss based on Kullback-Leibler (KL) divergence. Furthermore, the uncertainty in the feature domain is exploited by a novel perceptual loss; and we propose to calculate the adversarial loss from the gradient information in the SR stage for stable training performance and better visual quality. Experimental results on popular real-world datasets show that our proposed method has performed better than other unsupervised approaches. Qian Ning, Jingzhu Tang, Weisheng Dong, Xin Li 0005, Guangming Shi |
IJCAI | 4 |
| 2022 | AEDNet: Asynchronous Event Denoising with Spatial-Temporal Correlation among Irregular DataabstractDynamic Vision Sensor (DVS) is a compelling neuromorphic camera compared to conventional camera, but it suffers from fiercer noise. Due to the nature of irregular format and asynchronous readout, DVS data is always transformed into a regular tensor (e.g., 3D voxel or image) for deep learning method, which corrupts its own asynchronous properties. To maintain asynchronous, we establish an innovative asynchronous event denoise neural network, named AEDNet, which directly consumes the correlation of the irregular signal in spatial-temporal range without destroying its original structural property. Based on the property of continuation in temporal domain and discreteness in spatial domain, we decompose the DVS signal into two parts, i.e., temporal correlation and spatial affinity, and separately process these two parts. Our spatial feature embedding unit is a unique feature extraction module that extracts feature from event-level, which perfectly maintains its spatial-temporal correlation. To test effectiveness, we build a novel dataset named DVSCLEAN containing both simulated and real-world data. The experimental results of AEDNet achieve SOTA. Huachen Fang, Jinjian Wu, Leida Li, Junhui Hou, Weisheng Dong, Guangming Shi |
ACM Multimedia | 5 |
| 2022 | Bayesian based Re-parameterization for DNN Model PruningabstractFilter pruning, as an effective strategy to obtain efficient compact structures from over-parametric deep neural networks(DNN), has attracted a lot of attention. Previous pruning methods select channels for pruning by developing different criteria, yet little attention has been devoted to whether these criteria can represent correlations between channels. Meanwhile, most existing methods generally ignore the parameters being pruned and only perform additional training on the retained network to reduce accuracy loss. In this paper, we present a novel perspective of re-parametric pruning by Bayesian estimation. First, we estimate the probability distribution of different channels based on Bayesian estimation and indicate the importance of the channels by the discrepancy in the distribution before and after channel pruning. Second, to minimize the variation in distribution after pruning, we re-parameterize the pruned network based on the probability distribution to pursue optimal pruning. We evaluate our approach on popular datasets with some typical network architectures, and comprehensive experimental results validate that this method illustrates better performance compared to the state-of-the-art approaches. Xiaotong Lu, Teng Xi, Baopu Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 5 |
| 2022 | Learning for Motion Deblurring with Hybrid Frames and EventsabstractEvent camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events. Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 5 |
| 2022 | Robust Dynamic Background Modeling for Foreground EstimationabstractSeparating the background and foreground components from video frames is important to many tasks in computer vision and multimedia. As of today, robust principal component analysis (RPCA) has shown highly promising performance with the assumption that the background is low-rank and the foreground is sparse. However, existing RPCA-based methods have overlooked the uncertainty that some parts of the background (e.g., moving leaves in a dynamic background) or even the whole background (e.g., camera jittering) can be moving, which violates the low-rank assumption. To address this issue, we propose a novel enhanced RPCA framework (called ERPCA) by robustly modeling the dynamic background. Different from traditional RPCA framework, the background is decomposed into a low-rank component and a sparse component in the proposed ERPCA framework. Specifically, the sparse parts including foreground and dynamic parts of the background are modeled by Gaussian scale mixture (GSM) model. Moreover, those sparse components are further constrained by temporal consistency using nonzeromeans Gaussian models; the correspondences between sparse pixels in adjacent frames are explored by optical flow. Experimental results on 40 real videos demonstrate the superiority of our proposed method, with better average results than current state-of-the-art foreground estimation methods. Qian Ning, Weisheng Dong, Jinjian Wu, Guangming Shi, Xin Li 0005 |
VCIP | 3 |
| 2022 | Correlation filters based on spatial-temporal Gaussion scale mixture modelling for visual tracking
Guangming Shi, Weisheng Dong, Tianzhu Zhang 0001, Jinjian Wu, Xuemei Xie, Xin Li 0005 |
Neurocomputing | 3 |
| 2022 | Blind image quality assessment based on progressive multi-task learning
Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi |
Neurocomputing | 5 |
| 2022 | Bayesian Correlation Filter Learning With Gaussian Scale Mixture Model for Visual TrackingabstractCorrelation filters (CF), a popular tool for visual tracking, suffer from unwanted boundary effects due to the periodic assumption needed for FFT implementation. To address this issue, spatially regularized discriminative correlation filters (SRDCF) have been proposed by introducing a weighting matrix to the regularization term. However, the existing design of spatial weighting matrix is often heuristic and non-adaptive. Inspired by recent advances in joint discrimination and reliability learning for correlation tracking, we propose a principled Bayesian correlation filter learning method using Gaussian scale mixture (GSM) model. The key idea is to decompose each CF coefficient into the product of a positive scalar multiplier and a Gaussian random variable. Treating positive multipliers as weighting coefficients, GSM-based modeling of CFs leads to a spatially adaptive regularization strategy with improved capability of handling various appearance-related uncertainty factors (e.g., scale variation, out-of-plane rotation, and motion blur). Moreover, by imposing a sparse prior over the multipliers, we can jointly learn multipliers and CFs under a unified Bayesian estimation framework. Structured GSM model allows us to better exploit the spatial correlations among CFs and further improve the tracking performance. Experimental results on OTB-2013, OTB-2015, Temple Color-128, VOT-2016, and VOT-2017 show that our tracking method performs favorably when compared with current state-of-the-art methods. Guangming Shi, Tianzhu Zhang 0001, Weisheng Dong, Jinjian Wu, Xuemei Xie, Xin Li 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Blind Image Quality Index for Authentic Distortions With Local and Global Deep Feature AggregationabstractBlind image quality assessment (BIQA) for authentic distortions is still a great challenge, even in today’s deep learning era. It has been widely acknowledged that local and global features are both indispensable for IQA, which play complementary roles. While combining local and global features is straightforward in traditional handcrafted feature-based IQA metrics, it is not an easy task in the deep learning framework. This is mainly due to the fact that deep neural networks typically require input images with a fixed size. Current metrics either resize the image or use local patches as input, which are problematic in that they cannot integrate local and global aspects as well as their interactions to achieve comprehensive quality evaluation. Motivated by the above facts, this paper presents a new BIQA metric for authentic distortions by aggregating local and global deep features in a Vision-Transformer framework. In the proposed metric, selective local regions and global content are simultaneously input for complementary feature extraction, and the Vision-Transformer is employed to build the relationship between different local patches and image quality. Self-attention mechanism is further adopted to explore the interaction between local and global deep features, producing the final image quality score. Extensive experiments on five authentically distorted IQA databases demonstrate that the proposed metric outperforms the state-of-the-arts in terms of both prediction performance and generalization ability. Leida Li, Tianshu Song, Jinjian Wu, Weisheng Dong, Jiansheng Qian, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Spatiotemporal Representation Learning for Blind Video Quality AssessmentabstractBlind video quality assessment (BVQA) is of great importance for video-related applications, yet still challenging even in this deep learning era. The difficulty lies in the shortage of large-scale labeled data, thus making it hard to train a robust spatiotemporal encoder for BVQA. To relieve such difficulty, we first build a video dataset, which contains over 320K samples suffering from various compression and transmission artifacts. While manually annotating the dataset with subjective perception is much labor-intensive and time-consuming, we adopt reference-based VQA algorithms to weakly label the data automatically. We consider that single weak label is derived from single knowledge, which is deficient and incomplete for VQA. To alleviate the bias from single weak label (i.e., single knowledge) in the weakly labeled dataset, we propose HEterogeneous Knowledge Ensemble (HEKE) for spatiotemporal representation learning. Compared to learning from single knowledge, learning with HEKE is thought to achieve a lower infimum theoretically, and obtain richer representation. On the basis of the built dataset and the HEKE methodology, a feature encoder specific to BVQA is formed, and directly extract spatiotemporal representation from videos. Then, the video quality can be either acquired in a completely BVQA manner without ground truth, or via a finetuning-based regressor with labels. Extensive experiments on various VQA databases show that our BVQA model with the pretrained encoder achieves the state-of-the-art performance. More surprisingly, even trained on the synthetic data, our model still shows competitive performance on authentic databases. The data and source code will be available athttps://github.com/Sissuire/BVQA-HEKE. Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Generalizable No-Reference Image Quality Assessment via Deep Meta-LearningabstractRecently, researchers have shown great interest in using convolutional neural networks (CNNs) for no-reference image quality assessment (NR-IQA). Due to the lack of big training data, the efforts of existing metrics in optimizing CNN-based NR-IQA models remain limited. Furthermore, the diversity of distortions in images result in the generalization problem of NR-IQA models when trained with known distortions and tested on unseen distortions, which is an easy task for human. Hence, we propose a NR-IQA metric via deep meta-learning, which is highly generalizable in the face of unseen distortions. The fundamental idea is to learn the meta-knowledge shared by human when evaluating the quality of images with diversified distortions. Specifically, we define NR-IQA of different distortions as a series of tasks and propose a task selection strategy to build two task sets, which are characterized by synthetic to synthetic and synthetic to authentic distortions, respectively. Based on these two task sets, an optimization-based meta-learning is proposed to learn the generalized NR-IQA model, which can be directly used to evaluate the quality of images with unseen distortions. Extensive experiments demonstrate that our NR-IQA metric outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability. Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Contrastive Self-Supervised Pre-Training for Video Quality AssessmentabstractVideo quality assessment (VQA) task is an ongoing small sample learning problem due to the costly effort required for manual annotation. Since existing VQA datasets are of limited scale, prior research tries to leverage models pre-trained on ImageNet to mitigate this kind of shortage. Nonetheless, these well-trained models targeting on image classification task can be sub-optimal when applied on VQA data from a significantly different domain. In this paper, we make the first attempt to perform self-supervised pre-training for VQA task built upon contrastive learning method, targeting at exploiting the plentiful unlabeled video data to learn feature representation in a simple-yet-effective way. Specifically, we implement this idea by first generating distorted video samples with diverse distortion characteristics and visual contents based on the proposed distortion augmentation strategy. Afterwards, we conduct contrastive learning to capture quality-aware information by maximizing the agreement on feature representations of future frames and their corresponding predictions in the embedding space. In addition, we further introduce distortion prediction task as an additional learning objective to push the model towards discriminating different distortion categories of the input video. Solving these prediction tasks jointly with the contrastive learning not only provides stronger surrogate supervision signals, but also learns the shared knowledge among the prediction tasks. Extensive experiments demonstrate that our approach sets a new state-of-the-art in self-supervised learning for VQA task. Our results also underscore that the learned pre-trained model can significantly benefit the existing learning based VQA models. Source code is available at https://github.com/cpf0079/CSPT. Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2022 | Fine-Grained Image Quality Caption With Hierarchical Semantics DegradationabstractBlind image quality assessment (BIQA), which is capable of precisely and automatically estimating human perceived image quality with no pristine image for comparison, attracts extensive attention and is of wide applications. Recently, many existing BIQA methods commonly represent image quality with a quantitative value, which is inconsistent with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where humans are able to extract image contents with local to global semantics as they need. The mediocre quality value represents coarse or holistic image quality and fails to reflect degradation on hierarchical semantics. In this paper, to comply with human cognition, a novel quality caption model is inventively proposed to measure fine-grained image quality with hierarchical semantics degradation. Research on human visual system indicates there are hierarchy and reverse hierarchy correlations between hierarchical semantics. Meanwhile, empirical evidence shows that there are also bi-directional degradation dependencies between them. Thus, a novel bi-directional relationship-based network (BDRNet) is proposed for semantics degradation description, through adaptively exploring those correlations and degradation dependencies in a bi-directional manner. Extensive experiments demonstrate that our method outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability. Wen Yang 0008, Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 5 |
| 2022 | Video Quality Assessment With Serial Dependence ModelingabstractVideo quality assessment (VQA) is much more challenging than image quality assessment, due to the difficulty of modeling temporal influence among frames. Most of the existing VQA methods usually isolate each moment within the video (i.e., it neglects the sequential nature), leading to a large gap from the subjective perception. Recent research on neuroscience suggests a serially dependent perception (SDP) mechanism in the human visual system (HVS). Namely, the HVS tends to incorporate the recent past visual experience to predict the present perception. Inspired by the SDP, we suggest that the HVS prefers stable and continuous degradations in videos due to their predictability, and exhibits less tolerance to interrupted and unpredictable disturbances. Thus, we introduce a novel serial dependence modeling (SDM) framework for full-reference VQA in this paper. Firstly, the instantaneous degradation is measured on both the static appearance and motion information for each glimpse of scenes. Since motion plays an important role in videos, two types of structures are extracted for motion representation, namely, an explicit content-based 3D structure and an implicit feature-based 2D structure. Next, an assessment-directed long-short term memory (A-LSTM) is proposed to capture the serial dependence among instantaneous degradations. With the consideration of the perceptual effect from the previous moment on the current one, especially the effect from the perceptually worst moment, the serially dependent degradation is characterized. Finally, by mimicking the subjective rating for video-viewing, an attention-based quality decision procedure is presented to acquire the final video quality. Experimental results on publicly available VQA databases demonstrate that the proposed method maintains good consistency with the subjective perception. Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin |
IEEE Trans. Multim. | 5 |
| 2021 | Deep Gaussian Scale Mixture Prior for Spectral Compressive ImagingabstractIn coded aperture snapshot spectral imaging (CASSI) system, the real-world hyperspectral image (HSI) can be reconstructed from the captured compressive image in a snapshot. Model-based HSI reconstruction methods employed hand-crafted priors to solve the reconstruction problem, but most of which achieved limited success due to the poor representation capability of these hand-crafted priors. Deep learning based methods learning the mappings between the compressive images and the HSIs directly achieved much better results. Yet, it is nontrivial to design a powerful deep network heuristically for achieving satisfied results. In this paper, we propose a novel HSI reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Different from existing GSM models using hand-crafted scale priors (e.g., the Jeffrey’s prior), we propose to learn the scale prior through a deep convolutional neural network (DCNN). Furthermore, we also propose to estimate the local means of the GSM models by the DCNN. All the parameters of the MAP estimation algorithm and the DCNN parameters are jointly optimized through end-to-end training. Extensive experimental results on both synthetic and real datasets demonstrate that the proposed method outperforms existing state-of-the-art methods. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/DGSM-SCI.htm. Weisheng Dong, Xin Yuan 0002, Jinjian Wu, Guangming Shi |
CVPR | 2 |
| 2021 | Unsupervised Curriculum Domain Adaptation for No-Reference Video Quality AssessmentabstractDuring the last years, convolutional neural networks (C-NNs) have triumphed over video quality assessment (VQA) tasks. However, CNN-based approaches heavily rely on annotated data which are typically not available in VQA, leading to the difficulty of model generalization. Recent advances in domain adaptation technique makes it possible to adapt models trained on source data to unlabeled target data. However, due to the distortion diversity and content variation of the collected videos, the intrinsic subjectivity of VQA tasks hampers the adaptation performance. In this work, we propose a curriculum-style unsupervised domain adaptation to handle the cross-domain no-reference VQA problem. The proposed approach could be divided into two stages. In the first stage, we conduct an adaptation between source and target domains to predict the rating distribution for target samples, which can better reveal the subjective nature of VQA. From this adaptation, we split the data in target domain into confident and uncertain subdomains using the proposed uncertainty-based ranking function, through measuring their prediction confidences. In the second stage, by regarding samples in confident subdomain as the easy tasks in the curriculum, a fine-level adaptation is conducted between two subdomain-s to fine-tune the prediction model. Extensive experimental results on benchmark datasets highlight the superiority of the proposed method over the competing methods in both accuracy and speed. The source code is released at https://github.com/cpf0079/UCDA. Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
ICCV | 4 |
| 2021 | No-Reference Video Quality Assessment with Heterogeneous Knowledge EnsembleabstractBlind assessment of video quality is still challenging even in this deep learning era. The limited number of samples in existing databases is insufficient to learn a good feature extractor for video quality assessment (VQA), while manually labeling a larger database with subjective perception is very labor-intensive and time-consuming. To relieve such difficulty, we first collect 3589 high-quality video clips as the reference and build a large VQA dataset. The dataset contains more than 300K samples degraded by various distortion types due to compression and transmission error, and provides weak labels for each distorted sample with several full-reference VQA algorithms. To learn effective representation from the weakly labeled data, we alleviate the bias of single weak label (i.e., single knowledge) via learning from multiple heterogeneous knowledge. To this end, we propose a novel no-reference VQA (NR-VQA) method with HEterogeneous Knowledge Ensemble (HEKE). Comparing to learning from single knowledge, HEKE can theoretically reach a lower infimum, and learn richer representation due to the heterogeneity. Extensive experimental results show that the proposed HEKE outperforms existing NR-VQA methods, and achieves the state-of-the-art performance. The source code will be available at https://github.com/Sissuire/BVQA-HEKE. Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 4 |
| 2021 | Image Quality Caption with Attentive and Recurrent Semantic Attractor NetworkabstractIn this paper, a novel quality caption model is inventively developed to assess the image quality with hierarchical semantics. Existing image quality assessment (IQA) methods usually represent image quality with a quantitative value, resulting in inconsistency with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where hierarchical semantics are extracted. The mediocre quality value fails to reflect degradations on hierarchical semantics. Therefore, a new IQA framework is proposed to describe the quality for needs-oriented cognition. A novel quality caption procedure is firstly introduced, in which the quality is represented as patterns of activation distributed across the diverse degradations on hierarchical semantics. Then, an attentive and recurrent semantic attractor network (ARSANet) is designed to activate the distributed patterns for image quality description. Experiments demonstrate that our method achieves superior performance and is highly compliant with human cognition. Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi |
ACM Multimedia | 4 |
| 2021 | Uncertainty-Driven Loss for Single Image Super-ResolutionabstractIn low-level vision such as single image super-resolution (SISR), traditional MSE or L1 loss function treats every pixel equally with the assumption that the importance of all pixels is the same. However, it has been long recognized that texture and edge areas carry more important visual information than smooth areas in photographic images. How to achieve such spatial adaptation in a principled manner has been an open problem in both traditional model-based and modern learning-based approaches toward SISR. In this paper, we propose a new adaptive weighted loss for SISR to train deep networks focusing on challenging situations such as textured and edge pixels with high uncertainty. Specifically, we introduce variance estimation characterizing the uncertainty on a pixel-by-pixel basis into SISR solutions so the targeted pixels in a high-resolution image (mean) and their corresponding uncertainty (variance) can be learned simultaneously. Moreover, uncertainty estimation allows us to leverage conventional wisdom such as sparsity prior for regularizing SISR solutions. Ultimately, pixels with large certainty (e.g., texture and edge pixels) will be prioritized for SISR according to their importance to visual quality. For the first time, we demonstrate that such uncertainty-driven loss can achieve better results than MSE or L1 loss for a wide range of network architectures. Experimental results on three popular SISR networks show that our proposed uncertainty-driven loss has achieved better PSNR performance than traditional loss functions without any increased computation during testing. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UDL-SR.htm Qian Ning, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi |
NeurIPS | 2 |
| 2021 | Deep Maximum a Posterior Estimator for Video Denoising
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi |
Int. J. Comput. Vis. | 2 |
| 2021 | Toward blind joint demosaicing and denoising of raw color filter array data
Weisheng Dong, Guangming Shi, Zhonglong Zheng, Xin Li 0005 |
Neurocomputing | 3 |
| 2021 | Blind image quality prediction with hierarchical feature aggregation
Jinjian Wu, Wen Yang 0008, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin |
Inf. Sci. | 4 |
| 2021 | Robust subspace clustering network with dual-domain regularization
Guangming Shi, Xin Li 0005, Weisheng Dong, Jinjian Wu |
Pattern Recognit. Lett. | 5 |
| 2021 | Hybrid sparsity learning for image restoration: An iterative and trainable approach
Weisheng Dong, Guangming Shi, Shaoyuan Cheng, Xin Li 0005 |
Signal Process. | 2 |
| 2021 | Model-Guided Deep Hyperspectral Image Super-ResolutionabstractThe trade-off between spatial and spectral resolution is one of the fundamental issues in hyperspectral images (HSI). Given the challenges of directly acquiring high-resolution hyperspectral images (HR-HSI), a compromised solution is to fuse a pair of images: one has high-resolution (HR) in the spatial domain but low-resolution (LR) in spectral-domain and the other vice versa. Model-based image fusion methods including pan-sharpening aim at reconstructing HR-HSI by solving manually designed objective functions. However, such hand-crafted prior often leads to inevitable performance degradation due to a lack of end-to-end optimization. Although several deep learning-based methods have been proposed for hyperspectral pan-sharpening, HR-HSI related domain knowledge has not been fully exploited, leaving room for further improvement. In this paper, we propose an iterative Hyperspectral Image Super-Resolution (HSISR) algorithm based on a deep HSI denoiser to leverage both domain knowledge likelihood and deep image prior. By taking the observation matrix of HSI into account during the end-to-end optimization, we show how to unfold an iterative HSISR algorithm into a novel model-guided deep convolutional network (MoG-DCN). The representation of the observation matrix by subnetworks also allows the unfolded deep HSISR network to work with different HSI situations, which enhances the flexibility of MoG-DCN. Extensive experimental results are reported to demonstrate that the proposed MoG-DCN outperforms several leading HSISR methods in terms of both implementation cost and visual quality. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/MoG-DCN.htm. Weisheng Dong, Chen Zhou 0005, Jinjian Wu, Guangming Shi, Xin Li 0005 |
IEEE Trans. Image Process. | 1 |
| 2021 | Blind Image Quality Assessment With Active InferenceabstractBlind image quality assessment (BIQA) is a useful but challenging task. It is a promising idea to design BIQA methods by mimicking the working mechanism of human visual system (HVS). The internal generative mechanism (IGM) indicates that the HVS actively infers the primary content (i.e., meaningful information) of an image for better understanding. Inspired by that, this paper presents a novel BIQA metric by mimicking the active inference process of IGM. Firstly, an active inference module based on the generative adversarial network (GAN) is established to predict the primary content, in which the semantic similarity and the structural dissimilarity (i.e., semantic consistency and structural completeness) are both considered during the optimization. Then, the image quality is measured on the basis of its primary content. Generally, the image quality is highly related to three aspects, i.e., the scene information (content-dependency), the distortion type (distortion-dependency), and the content degradation (degradation-dependency). According to the correlation between the distorted image and its primary content, the three aspects are analyzed and calculated respectively with a multi-stream convolutional neural network (CNN) based quality evaluator. As a result, with the help of the primary content obtained from the active inference and the comprehensive quality degradation measurement from the multi-stream CNN, our method achieves competitive performance on five popular IQA databases. Especially in cross-database evaluations, our method achieves significant improvements. Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie, Guangming Shi, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2021 | Probabilistic Undirected Graph Based Denoising Method for Dynamic Vision SensorabstractDynamic Vision Sensor (DVS) is a new type of neuromorphic event-based sensor, which has an innate advantage in capturing fast-moving objects. Due to the interference of DVS hardware itself and many external factors, noise is unavoidable in the output of DVS. Different from frame/image with structural data, the output of DVS is in the form of address-event representation (AER), which means that the traditional denoising methods cannot be used for the output (i.e., event stream) of the DVS. In this paper, we propose a novel event stream denoising method based on probabilistic undirected graph model (PUGM). The motion of objects always shows a certain regularity/trajectory in space and time, which reflects the spatio-temporal correlation between effective events in the stream. Meanwhile, the event stream of DVS is composed by the effective events and random noise. Thus, a probabilistic undirected graph model is constructed to describe such priori knowledge (i.e., spatio-temporal correlation). The undirected graph model is factorized into the product of the cliques energy function, and the energy function is defined to obtain the complete expression of the joint probability distribution. Better denoising effect means a higher probability (lower energy), which means the denoising problem can be transfered into energy optimization problem. Thus, the iterated conditional modes (ICM) algorithm is used to optimize the model to remove the noise. Experimental results on denoising show that the proposed algorithm can effectively remove noise events. Moreover, with the preprocessing of the proposed algorithm, the recognition accuracy on AER data can be remarkably promoted. Jinjian Wu, Chuanwei Ma, Leida Li, Weisheng Dong, Guangming Shi |
IEEE Trans. Multim. | 4 |
| 2020 | Spatial-Temporal Gaussian Scale Mixture Modeling for Foreground EstimationabstractSubtracting the backgrounds from the video frames is an important step for many video analysis applications. Assuming that the backgrounds are low-rank and the foregrounds are sparse, the robust principle component analysis (RPCA)-based methods have shown promising results. However, the RPCA-based methods suffered from the scale issue, i.e., the ℓ1-sparsity regularizer fails to model the varying sparsity of the moving objects. While several efforts have been made to address this issue with advanced sparse models, previous methods cannot fully exploit the spatial-temporal correlations among the foregrounds. In this paper, we proposed a novel spatial-temporal Gaussian scale mixture (STGSM) model for foreground estimation. In the proposed STGSM model, a temporal consistent constraint is imposed over the estimated foregrounds through nonzero-means Gaussian models. Specifically, the estimates of the foregrounds obtained in the previous frame are used as the prior for these of the current frame, and nonzero means Gaussian scale mixture models (GSM) are developed. To better characterize the temporal correlations, the optical flow has been used to model the correspondences between foreground pixels in adjacent frames. The spatial correlations have also been exploited by considering that local correlated pixels should be characterized by the same STGSM model, leading to further performance improvements. Experimental results on real video datasets show that the proposed method performs comparably or even better than current state-of-the-art background subtraction methods. Qian Ning, Weisheng Dong, Jinjian Wu, Jie Lin 0008, Guangming Shi |
AAAI | 2 |
| 2020 | MetaIQA: Deep Meta-Learning for No-Reference Image Quality AssessmentabstractRecently, increasing interest has been drawn in exploiting deep convolutional neural networks (DCNNs) for no-reference image quality assessment (NR-IQA). Despite of the notable success achieved, there is a broad consensus that training DCNNs heavily relies on massive annotated data. Unfortunately, IQA is a typical small sample problem. Therefore, most of the existing DCNN-based IQA metrics operate based on pre-trained networks. However, these pre-trained networks are not designed for IQA task, leading to generalization problem when evaluating different types of distortions. With this motivation, this paper presents a no-reference IQA metric based on deep meta-learning. The underlying idea is to learn the meta-knowledge shared by human when evaluating the quality of images with various distortions, which can then be adapted to unknown distortions easily. Specifically, we first collect a number of NR-IQA tasks for different distortions. Then meta-learning is adopted to learn the prior knowledge shared by diversified distortions. Finally, the quality prior model is fine-tuned on a target NR-IQA task for quickly obtaining the quality model. Extensive experiments demonstrate that the proposed metric outperforms the state-of-the-arts by a large margin. Furthermore, the meta-model learned from synthetic distortions can also be easily generalized to authentic distortions, which is highly desired in real-world applications of IQA metrics. Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
CVPR | 4 |
| 2020 | Active Inference of GAN for No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) is a challenging task. It is a promising idea to design NR-IQA algorithms by mimicking how human visual system (HVS) works. The internal generative mechanism (IGM) indicates that HVS actively infers the primary content of an image for better understanding. Inspired by that, a novel NR-IQA method with active inference is proposed in this paper. First, a generative adversarial network (GAN) is proposed to predict the primary content of a distorted image, in which two IGM-inspired constraints are considered during the optimization. Next, based on the correlation between the distorted image and its primary content, different degradations (i.e., the content/distortion-/structure-dependency degradation) are measured simultaneously with a multi-stream convolutional neural network (CNN) for NR-IQA. Benefit from the primary content obtained from GAN and the multiple degradations measurement of CNN, our method achieves the state-of-the-art on five public IQA databases. Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie |
ICME | 4 |
| 2020 | Beyond Network Pruning: a Joint Search-and-Training ApproachabstractNetwork pruning has been proposed as a remedy for alleviating the over-parameterization problem of deep neural networks. However, its value has been recently challenged especially from the perspective of neural architecture search (NAS). We challenge the conventional wisdom of pruning-after-training by proposing a joint search-and-training approach that directly learns a compact network from the scratch. By treating pruning as a search strategy, we present two new insights in this paper: 1) it is possible to expand the search space of networking pruning by associating each filter with a learnable weight; 2) joint search-and-training can be conducted iteratively to maximize the learning efficiency. More specifically, we propose a coarse-to-fine tuning strategy to iteratively sample and update compact sub-network to approximate the target network. The weights associated with network filters will be accordingly updated by joint search-and-training to reflect learned knowledge in NAS space. Moreover, we introduce strategies of random perturbation (inspired by Monte Carlo) and flexible thresholding (inspired by Reinforcement Learning) to adjust the weight and size of each layer. Extensive experiments on ResNet and VGGNet demonstrate the superior performance of our proposed method on popular datasets including CIFAR10, CIFAR100 and ImageNet. Xiaotong Lu, Weisheng Dong, Xin Li 0005, Guangming Shi |
IJCAI | 3 |
| 2020 | End-to-End Blind Image Quality Prediction With Cascaded Deep Neural NetworkabstractThe deep convolutional neural network (CNN) has achieved great success in image recognition. Many image quality assessment (IQA) methods directly use recognition-oriented CNN for quality prediction. However, the properties of IQA task is different from image recognition task. Image recognition should be sensitive to visual content and robust to distortion, while IQA should be sensitive to both distortion and visual content. In this paper, an IQA-oriented CNN method is developed for blind IQA (BIQA), which can efficiently represent the quality degradation. CNN is large-data driven, while the sizes of existing IQA databases are too small for CNN optimization. Thus, a large IQA dataset is firstly established, which includes more than one million distorted images (each image is assigned with a quality score as its substitute of Mean Opinion Score (MOS), abbreviated as pseudo-MOS). Next, inspired by the hierarchical perception mechanism (from local structure to global semantics) in human visual system, a novel IQA-orientated CNN method is designed, in which the hierarchical degradation is considered. Finally, by jointly optimizing the multilevel feature extraction, hierarchical degradation concatenation (HDC) and quality prediction in an end-to-end framework, the Cascaded CNN with HDC (named as CaHDC) is introduced. Experiments on the benchmark IQA databases demonstrate the superiority of CaHDC compared with existing BIQA methods. Meanwhile, the CaHDC (with about 0.73M parameters) is lightweight comparing to other CNN-based BIQA models, which can be easily realized in the microprocessing system. The dataset and source code of the proposed method are available at https://web.xidian.edu.cn/wjj/paper.html. Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2019 | End-to-End Blind Image Quality Assessment with Cascaded Deep FeaturesabstractThe convolutional neural network (CNN) has achieved great success in many visual tasks. However, it has limited progress on image quality assessment (IQA) due to the lacking of IQA-oriented CNN framework which can efficiently represent the hierarchical quality degradation. In this paper, inspired by the hierarchical perception mechanism (from local structure to global semantics) in the human visual system, we design an end-to-end cascaded CNN framework for blind IQA (BIQA), in which multilevel features are extracted and concatenated to represent the hierarchical quality degradation. By jointly optimizing the feature extraction, hierarchical degradation integration, and quality prediction in an end-to-end manner, the novel cascaded CNN with hierarchical feature integration (CaHFI) for BIQA is designed. Experimental results on five benchmark IQA databases demonstrate that the proposed CaHFI achieves the state-of-the-art. And experiments on cross-database evaluation further prove the high generalization ability of the proposed CaHFI. Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi |
ICME | 4 |
| 2019 | SISRSet: Single image super-resolution subjective evaluation test and objective quality assessment
Guangming Shi, Wenfei Wan, Jinjian Wu, Xuemei Xie, Weisheng Dong, Hong Ren Wu |
Neurocomputing | 5 |
| 2019 | No-reference image quality assessment with visual pattern degradation
Jinjian Wu, Man Zhang 0007, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin |
Inf. Sci. | 4 |
| 2019 | Blind image quality assessment with hierarchy: Degradation from local structure to deep semantics
Jinjian Wu, Jichen Zeng, Weisheng Dong, Guangming Shi, Weisi Lin |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Denoising Prior Driven Deep Neural Network for Image RestorationabstractDeep neural networks (DNNs) have shown very promising results for various image restoration (IR) tasks. However, the design of network architectures remains a major challenging for achieving further improvements. While most existing DNN-based methods solve the IR problems by directly mapping low quality images to desirable high-quality images, the observation models characterizing the image degradation processes have been largely ignored. In this paper, we first propose a denoising-based IR algorithm, whose iterative steps can be computed efficiently. Then, the iterative process is unfolded into a deep neural network, which is composed of multiple denoisers modules interleaved with back-projection (BP) modules that ensure the observation consistencies. A convolutional neural network (CNN) based denoiser that can exploit the multi-scale redundancies of natural images is proposed. As such, the proposed network not only exploits the powerful denoising ability of DNNs, but also leverages the prior of the observation model. Through end-to-end training, both the denoisers and the BP modules can be jointly optimized. Experimental results on several IR tasks, e.g., image denoisig, super-resolution and deblurring show that the proposed method can lead to very competitive and often state-of-the-art results on several IR tasks, including image denoising, deblurring, and super-resolution. Weisheng Dong, Wotao Yin, Guangming Shi, Xiaotong Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Multi-layer discriminative dictionary learning with locality constraint for image classification
Jianqiang Song, Xuemei Xie, Guangming Shi, Weisheng Dong |
Pattern Recognit. | 4 |
| 2019 | Image Caption Generation with Part of Speech Guidance
Xinwei He 0001, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang 0001, Weisheng Dong |
Pattern Recognit. Lett. | 6 |
| 2019 | Quality Assessment for Video With Degradation Along Salient TrajectoriesabstractWith the rapid growth of digital video through the Internet, a reliable objective video-quality assessment (VQA) algorithm is in great demand for video management. Motion information plays a dominant role for video perception, and the human visual system (HVS) is able to track moving objects effectively with eye movement. Moreover, the middle temporal area of the brain is selective for moving objects with particular velocities. In other words, visual contents that are along the motion trajectories will automatically attract our attention for dedicated processing. Inspired by the motion-related process in the HVS, we suggest analyzing the degradation along attended motion trajectories for VQA. The characteristic of motion velocity along each trajectory is analyzed for temporal quality measurement. Meanwhile, visual information along each trajectory is extracted for joint spatial-temporal quality measurement. Finally, considering the spatial-quality degradation from each frame, a novel full-reference assessor along salient trajectories (FAST) for VQA (which combines the spatial, temporal, and joint spatial-temporal quality degradations) is introduced. Experimental results on five publicly available VQA databases demonstrate that the proposed FAST VQA model performs consistently with the subjective perception. The source code of the proposed method is available at http://web.xidian.edu.cn/wjj/paper.html. Jinjian Wu, Yongxu Liu 0001, Weisheng Dong, Guangming Shi, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2018 | Super-Resolution Quality Assessment: Subjective Evaluation Database and Quality Index Based on Perceptual Structure MeasurementabstractWith the outstanding performance of deep learning based single image super-resolution (SISR) methods, the traditional SISR evaluation metrics (e.g., PSNR and SSIM, which measure the per-pixel differences and simple structure similarities respectively) are facing great challenges. When assessing SISR algorithms, they generally are hardly consistent with the human visual system (HVS). According to the psychological studies, the HVS presents different sensitivities to the plain, edge and texture regions, which are difficult to be accurately identified and measured with the existing quality indexes, especially for SR images. To deal with this problem, we firstly build a SISR subjective assessment database including several major deep learning based SR methods. Then we propose a more accurate perception structure measurement and use their similarity comparisons to evaluate the SR algorithms. Experimental results on the databases demonstrate that the proposed method performs well consistent with the human visual perception. Wenfei Wan, Jinjian Wu, Guangming Shi, Weisheng Dong |
ICME | 5 |
| 2018 | Lightweight Deep Residue Learning for Joint Color Image Demosaicking and DenoisingabstractColor demosaicking and image denoising each plays an important role in digital cameras. Conventional model-based methods often fail around the areas of strong textures and produce disturbing visual artifacts such as aliasing and zippering. Recently developed deep learning based methods were capable of obtaining images of better qualities though at the price of high computational cost, which make them not suitable for real-time applications. In this paper, we propose a lightweight convolutional neural network for joint demosaicking and denoising (JDD) problem with the following salient features. First, the densely connected network is trained in an end-to-end manner to learn the mapping from the noisy low-resolution space (CFA image) to the clean high-resolution space (color image). Second, the concept of deep residue learning and aggregated residual transformations are extended from image denoising and classification to JDD supporting more efficient training. Third, the design of our end-to-end network architecture is inspired by a rigorous analysis of JDD using sparsity models. Experimental results conducted for both demosaicking-only and JDD tasks have shown that the proposed method performs much better than existing state-of-the-art methods (i.e., higher visual quality, smaller training set and lower computational cost). Weisheng Dong, Guangming Shi, Xin Li 0005 |
ICPR | 3 |
| 2018 | Exploiting class-wise coding coefficients: Learning a discriminative dictionary for pattern classification
Jianqiang Song, Xuemei Xie, Guangming Shi, Weisheng Dong |
Neurocomputing | 4 |
| 2018 | Image Super-Resolution With Parametric Sparse Model LearningabstractRecovering a high-resolution (HR) image from its low-resolution (LR) version is an ill-posed inverse problem. Learning accurate prior of HR images is of great importance to solve this inverse problem. Existing super-resolution (SR) methods either learn a non-parametric image prior from training data (a large set of LR/HR patch pairs) or estimate a parametric prior from the LR image analytically. Both methods have their limitations: the former lacks flexibility when dealing with different SR settings; while the latter often fails to adapt to spatially varying image structures. In this paper, we propose to take a hybrid approach toward image SR by combining those two lines of ideas - that is, a parametric sparse prior of HR images is learned from the training set as well as the input LR image. By exploiting the strengths of both worlds, we can more accurately recover the sparse codes and therefore HR image patches than conventional sparse coding approaches. Experimental results show that the proposed hybrid SR method significantly outperforms existing model-based SR methods and is highly competitive to current state-of-the-art learning-based SR methods in terms of both subjective and objective image qualities. Weisheng Dong, Xuemei Xie, Guangming Shi, Jinjian Wu, Xin Li 0005 |
IEEE Trans. Image Process. | 2 |
| 2018 | Robust Foreground Estimation via Structured Gaussian Scale Mixture ModelingabstractRecovering the background and foreground parts from video frames has important applications in video surveillance. Under the assumption that the background parts are stationary and the foreground are sparse, most of existing methods are based on the framework of robust principal component analysis (RPCA), i.e., modeling the background and foreground parts as a low-rank and sparse matrices, respectively. However, in realistic complex scenarios, the conventional norm sparse regularizer often fails to well characterize the varying sparsity of the foreground components. How to select the sparsity regularizer parameters adaptively according to the local statistics is critical to the success of the RPCA framework for background subtraction task. In this paper, we propose to model the sparse component with a Gaussian scale mixture (GSM) model. Compared with the conventional norm, the GSM-based sparse model has the advantages of jointly estimating the variances of the sparse coefficients (and hence the regularization parameters) and the unknown sparse coefficients, leading to significant estimation accuracy improvements. Moreover, considering that the foreground parts are highly structured, a structured extension of the GSM model is further developed. Specifically, the input frame is divided into many homogeneous regions using superpixel segmentation. By characterizing the set of sparse coefficients in each homogeneous region with the same GSM prior, the local dependencies among the sparse coefficients can be effectively exploited, leading to further improvements for background subtraction. Experimental results on several challenging scenarios show that the proposed method performs much better than most of existing background subtraction methods in terms of both performance and speed. Guangming Shi, Weisheng Dong, Jinjian Wu, Xuemei Xie |
IEEE Trans. Image Process. | 3 |
| 2018 | Blind Quality Index for Multiply Distorted Images Using Biorder Structure Degradation and Nonlocal StatisticsabstractIn the past decade, extensive image quality metrics have been proposed. The majority of them are tailored for the images that contain a specific type of distortion. However, in practice, the images are usually degraded by different types of distortions simultaneously. This poses great challenges to the existing quality metrics. Motivated by this, this paper proposes a no-reference quality index for the multiply distorted images using the biorder structure degradation and the nonlocal statistics. The design philosophy is inspired by the fact that the human visual system (HVS) is highly sensitive to the degradations of both the spatial contrast and the spatial distribution, which are prone to be changed by the joint effects of the multiple distortions. Specifically, the multiresolution representation of the image is first built by downsampling to simulate the hierarchical property of the HVS. Then, the structure degradation is calculated to measure the spatial contrast. Considering the fact that the human visual cortex has the separate mechanisms to perceive the first- and second-order structures, dubbed biorder structures, the degradations of biorder structures are calculated to account for the spatial contrast, producing the first group of the quality-aware features. Furthermore, the nonlocal self-similarity statistics is calculated to measure the spatial distribution, producing the second group of features. Finally, all the features are fed into the random forest regression model to learn the quality model for the multiply distorted images. Extensive experimental results conducted on the three public databases demonstrate the superiority of the proposed metric to the state-of-the-art metrics. Moreover, the proposed metric is also advantageous over the existing metrics in terms of the generalization ability. Yu Zhou 0009, Leida Li, Jinjian Wu, Ke Gu 0001, Weisheng Dong, Guangming Shi |
IEEE Trans. Multim. | 5 |
| 2017 | Bag-of-words feature representation for blind image quality assessment with local quantized pattern
Xuemei Xie, Yazhong Zhang, Jinjian Wu, Guangming Shi, Weisheng Dong |
Neurocomputing | 5 |
| 2017 | Mixed Noise Removal via Laplacian Scale Mixture Modeling and Nonlocal Low-Rank ApproximationabstractRecovering the image corrupted by additive white Gaussian noise (AWGN) and impulse noise is a challenging problem due to its difficulties in an accurate modeling of the distributions of the mixture noise. Many efforts have been made to first detect the locations of the impulse noise and then recover the clean image with image in painting techniques from an incomplete image corrupted by AWGN. However, it is quite challenging to accurately detect the locations of the impulse noise when the mixture noise is strong. In this paper, we propose an effective mixture noise removal method based on Laplacian scale mixture (LSM) modeling and nonlocal low-rank regularization. The impulse noise is modeled with LSM distributions, and both the hidden scale parameters and the impulse noise are jointly estimated to adaptively characterize the real noise. To exploit the nonlocal self-similarity and low-rank nature of natural image, a nonlocal low-rank regularization is adopted to regularize the denoising process. Experimental results on synthetic noisy images show that the proposed method outperforms existing mixture noise removal methods. Weisheng Dong, Xuemei Xie, Guangming Shi, Xiang Bai |
IEEE Trans. Image Process. | 2 |
| 2017 | Enhanced Just Noticeable Difference Model for Images With Pattern ComplexityabstractThe just noticeable difference (JND) in an image, which reveals the visibility limitation of the human visual system (HVS), is widely used for visual redundancy estimation in signal processing. To determine the JND threshold with the current schemes, the spatial masking effect is estimated as the contrast masking, and this cannot accurately account for the complicated interaction among visual contents. Research on cognitive science indicates that the HVS is highly adapted to extract the repeated patterns for visual content representation. Inspired by this, we formulate the pattern complexity as another factor to determine the total masking effect: the interaction is relatively straightforward with a limited masking effect in a regular pattern, and is complicated with a strong masking effect in an irregular pattern. From the orientation selectivity mechanism in the primary visual cortex, the response of each local receptive field can be considered as a pattern; therefore, in this paper, the orientation that each pixel presents is regarded as the fundamental element of a pattern, and the pattern complexity is calculated as the diversity of the orientation in a local region. Finally, considering both pattern complexity and luminance contrast, a novel spatial masking estimation function is deduced, and an improved JND estimation model is built. Experimental results on comparing with the latest JND models demonstrate the effectiveness of the proposed model, which performs highly consistent with the human perception. The source code of the proposed model is publicly available at http://web.xidian.edu.cn/wjj/en/index.html. Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 3 |
| 2017 | Color-Guided Depth Recovery via Joint Local Structural and Nonlocal Low-Rank RegularizationabstractHigh-quality depth recovery from RGB-D data has received increasingly more attention in recent years due to their wide applications from depth-based image rendering to three-dimensional imaging and video. Sharp contrast between high-quality color images and low-quality depth maps presents severe challenges to the development of color-guided depth recovery techniques. Previous works have emphasized either locally varying characteristics of color-depth dependence or nonlocal similarities around the discontinuities of the scene geometry. Therefore, it is desirable to exploit both local and nonlocal structural constraints for optimizing the performance of color-guided depth recovery. In this work, we propose a unified variational approach via joint local and nonlocal regularization. The local regularization term consists of two complementary parts-one characterizing the color-depth dependence in the gradient domain and the other in the spatial domain; nonlocal regularization involves a low-rank constraint suitable for large-scale depth discontinuities. Extensive experimental results are reported to show that our approach outperforms several existing state-of-the-art depth recovery methods on both synthetic and real-world data sets. Weisheng Dong, Guangming Shi, Xin Li 0005, Kefan Peng, Jinjian Wu, Zhenhua Guo 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Learning Parametric Sparse Models for Image Super-ResolutionabstractLearning accurate prior knowledge of natural images is of great importance for single image super-resolution (SR). Existing SR methods either learn the prior from the low/high-resolution patch pairs or estimate the prior models from the input low-resolution (LR) image. Specifically, high-frequency details are learned in the former methods. Though effective, they are heuristic and have limitations in dealing with blurred LR images; while the latter suffers from the limitations of frequency aliasing. In this paper, we propose to combine those two lines of ideas for image super-resolution. More specifically, the parametric sparse prior of the desirable high-resolution (HR) image patches are learned from both the input low-resolution (LR) image and a training image dataset. With the learned sparse priors, the sparse codes and thus the HR image patches can be accurately recovered by solving a sparse coding problem. Experimental results show that the proposed SR method outperforms existing state-of-the-art methods in terms of both subjective and objective image qualities. Weisheng Dong, Xuemei Xie, Guangming Shi, Xin Li 0005, Donglai Xu |
NIPS | 2 |
| 2016 | 3D magnetic resonance image denoising using low-rank tensor approximation
Weisheng Dong |
Neurocomputing | 2 |
| 2016 | Hyperspectral Image Super-Resolution via Non-Negative Structured Sparse RepresentationabstractHyperspectral imaging has many applications from agriculture and astronomy to surveillance and mineralogy. However, it is often challenging to obtain high-resolution (HR) hyperspectral images using existing hyperspectral imaging techniques due to various hardware limitations. In this paper, we propose a new hyperspectral image super-resolution method from a low-resolution (LR) image and a HR reference image of the same scene. The estimation of the HR hyperspectral image is formulated as a joint estimation of the hyperspectral dictionary and the sparse codes based on the prior knowledge of the spatial-spectral sparsity of the hyperspectral image. The hyperspectral dictionary representing prototype reflectance spectra vectors of the scene is first learned from the input LR image. Specifically, an efficient non-negative dictionary learning algorithm using the block-coordinate descent optimization technique is proposed. Then, the sparse codes of the desired HR hyperspectral image with respect to learned hyperspectral basis are estimated from the pair of LR and HR reference images. To improve the accuracy of non-negative sparse coding, a clustering-based structured sparse coding method is proposed to exploit the spatial correlation among the learned sparse codes. The experimental results on both public datasets and real LR hypspectral images suggest that the proposed method substantially outperforms several existing HR hyperspectral image recovery techniques in the literature in terms of both objective quality metrics and computational efficiency. Weisheng Dong, Fazuo Fu, Guangming Shi, Xun Cao, Jinjian Wu, Xin Li 0005 |
IEEE Trans. Image Process. | 1 |
| 2015 | Low-Rank Tensor Approximation with Laplacian Scale Mixture Modeling for Multiframe Image DenoisingabstractPatch-based low-rank models have shown effective in exploiting spatial redundancy of natural images especially for the application of image denoising. However, two-dimensional low-rank model can not fully exploit the spatio-temporal correlation in larger data sets such as multispectral images and 3D MRIs. In this work, we propose a novel low-rank tensor approximation framework with Laplacian Scale Mixture (LSM) modeling for multi-frame image denoising. First, similar 3D patches are grouped to form a tensor of d-order and high-order Singular Value Decomposition (HOSVD) is applied to the grouped tensor. Then the task of multiframe image denoising is formulated as a Maximum A Posterior (MAP) estimation problem with the LSM prior for tensor coefficients. Both unknown sparse coefficients and hidden LSM parameters can be efficiently estimated by the method of alternating optimization. Specifically, we have derived closed-form solutions for both subproblems. Experimental results on spectral and dynamic MRI images show that the proposed algorithm can better preserve the sharpness of important image structures and outperform several existing state-of-the-art multiframe denoising methods (e.g., BM4D and tensor dictionary learning). Weisheng Dong, Guangming Shi, Xin Li 0005, Yi Ma 0001 |
ICCV | 1 |
| 2015 | Learning Parametric Distributions for Image Super-Resolution: Where Patch Matching Meets Sparse CodingabstractExisting approaches toward Image super-resolution (SR) is often either data-driven (e.g., based on internet-scale matching and web image retrieval) or model-based (e.g., formulated as an Maximizing a Posterior estimation problem). The former is conceptually simple yet heuristic, while the latter is constrained by the fundamental limit of frequency aliasing. In this paper, we propose to develop a hybrid approach toward SR by combining those two lines of ideas. More specifically, the parameters underlying sparse distributions of desirable HR image patches are learned from a pair of LR image and retrieved HR images. Our hybrid approach can be interpreted as the first attempt of reconciling the difference between parametric and nonparametric models for low-level vision tasks. Experimental results show that the proposed hybrid SR method performs much better than existing state-of-the-art methods in terms of both subjective and objective image qualities. Weisheng Dong, Guangming Shi, Xuemei Xie |
ICCV | 2 |
| 2015 | Image Restoration via Simultaneous Sparse Coding: Where Structured Sparsity Meets Gaussian Scale Mixture
Weisheng Dong, Guangming Shi, Yi Ma 0001, Xin Li 0005 |
Int. J. Comput. Vis. | 1 |
| 2015 | Visual Orientation Selectivity Based Structure DescriptionabstractThe human visual system is highly adaptive to extract structure information for scene perception, and structure character is widely used in perception-oriented image processing works. However, the existing structure descriptors mainly describe the luminance contrast of a local region, but cannot effectively represent the spatial correlation of structure. In this paper, we introduce a novel structure descriptor according to the orientation selectivity mechanism in the primary visual cortex. Research on cognitive neuroscience indicate that the arrangement of excitatory and inhibitory cortex cells arise orientation selectivity in a local receptive field, within which the primary visual cortex performs visual information extraction for scene understanding. Inspired by the orientation selectivity mechanism, we compute the correlations among pixels in a local region based on the similarities of their preferred orientation. By imitating the arrangement of the excitatory/inhibitory cells, the correlations between a central pixel and its local neighbors are binarized, and the spatial correlation is represented with a set of binary values, which is named the orientation selectivity-based pattern. Then, taking both the gradient magnitude and the orientation selectivity-based pattern into account, a rotation invariant structure descriptor is introduced. The proposed structure descriptor is applied in texture classification and reduced reference image quality assessment, as two different application domains to verify its generality and robustness. Experimental results demonstrate that the orientation selectivity-based structure descriptor is robust to disturbance, and can effectively represent the structure degradation caused by different types of distortion. Jinjian Wu, Weisi Lin, Guangming Shi, Yazhong Zhang, Weisheng Dong, Zhibo Chen 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | Sparsity fine tuning in wavelet domain with application to compressive image reconstructionabstractIn compressive sensing, wavelet space is widely used to generate sparse signal (image signal in particular) representations. In this work, we propose a novel approach of statistical context modeling to increase the level of sparsity of wavelet image representations. It is shown, contrary to a widely held assumption, that high-frequency wavelet coefficients have non-zero mean distributions if conditioned on local image structures. Removing this bias can make wavelet image representations sparser, i.e., having a greater number of zero and close-to-zero coefficients. The resulting unbiased probability models can significantly improve the performance of existing wavelet-based compressive image reconstruction methods in both PSNR and visual quality. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
ICASSP | 1 |
| 2014 | Image restoration via Bayesian structured sparse codingabstractIn this work, we propose a Bayesian structured sparse coding (BSSC) framework containing a nonlocal extension of Gaussian scale mixture (GSM) model by exploiting structured sparsity. It is shown that the variances of sparse coefficients (the field of Gaussian scalars) - if treated as a latent variable - can besparse coefficients jointly estimated along with the unknown sparse coefficients via the the method of alternative optimization. When applied to image restoration, BSSC leads to closed-form solutions involving iterative shrinkage/filtering and therefore admits computationally efficient implementation. Our experimental results have shown that BSSC-based image restoration often delivers reconstructed images with higher subjective/objective qualities than other competing approaches including IDD-BM3D and NCSR. Weisheng Dong, Xin Li 0005, Yi Ma 0001, Guangming Shi |
ICIP | 1 |
| 2014 | Nonlocal Sparse and Low-Rank Regularization for Optical Flow EstimationabstractDesigning an appropriate regularizer is of great importance for accurate optical flow estimation. Recent works exploiting the nonlocal similarity and the sparsity of the motion field have led to promising flow estimation results. In this paper, we propose to unify these two powerful priors. To this end, we propose an effective flow regularization technique based on joint low-rank and sparse matrix recovery. By grouping similar flow patches into clusters, we effectively regularize the motion field by decomposing each set of similar flow patches into a low-rank component and a sparse component. For better enforcing the low-rank property, instead of using the convex nuclear norm, we use the log det(·) function as the surrogate of rank, which can also be efficiently minimized by iterative singular value thresholding. Experimental results on the Middlebury benchmark show that the performance of the proposed nonlocal sparse and low-rank regularization method is higher than (or comparable to) those of previous approaches that harness these same priors, and is competitive to current state-of-the-art methods. Weisheng Dong, Guangming Shi, Xiaocheng Hu, Yi Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Compressive Sensing via Nonlocal Low-Rank RegularizationabstractSparsity has been widely exploited for exact reconstruction of a signal from a small number of random measurements. Recent advances have suggested that structured or group sparsity often leads to more powerful signal reconstruction techniques in various compressed sensing (CS) studies. In this paper, we propose a nonlocal low-rank regularization (NLR) approach toward exploiting structured sparsity and explore its application into CS of both photographic and MRI images. We also propose the use of a nonconvex log det ( X) as a smooth surrogate function for the rank instead of the convex nuclear norm and justify the benefit of such a strategy using extensive experiments. To further improve the computational efficiency of the proposed algorithm, we have developed a fast implementation using the alternative direction multiplier method technique. Experimental results have shown that the proposed NLR-CS algorithm can significantly outperform existing state-of-the-art CS techniques for image recovery. Weisheng Dong, Guangming Shi, Xin Li 0005, Yi Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Sparsity Fine Tuning in Wavelet Domain With Application to Compressive Image ReconstructionabstractIn compressive sensing, wavelet space is widely used to generate sparse signal (image signal in particular) representations. In this paper, we propose a novel approach of statistical context modeling to increase the level of sparsity of wavelet image representations. It is shown, contrary to a widely held assumption, that high-frequency wavelet coefficients have nonzero mean distributions if conditioned on local image structures. Removing this bias can make wavelet image representations sparser, i.e., having a greater number of zero and closeto-zero coefficients. The resulting unbiased probability models can significantly improve the performance of existing wavelet-based compressive image reconstruction methods in both PSNR and visual quality. An efficient algorithm is presented to solve the compressive image recovery (CIR) problem using the refined models. Experimental results on both simulated compressive sensing (CS) image data and real CS image data show that the new CIR method significantly outperforms existing CIR methods in both PSNR and visual quality. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Image Process. | 1 |
| 2013 | SAR Image Quality Assessment Based on SSIM Using Textural FeatureabstractThe Synthetic Aperture Radar (SAR) image quality assessment (IQA) can provide a measurement for SAR jamming effect, which is helpful to improve the jamming pattern. A texture-based SSIM (TSSIM) algorithm is proposed, because of the fact that SAR images have much texture and SSIM algorithm has good performance in optical image quality assessment. TSSIM combines the image gray intensity with textural features to measure the image structural information. The Gray Level Co-occurrence Matrix (GLCM) is adopted to extract the textural features of SAR image. The angle second moment feature is proved the best compared with the other textural features. The results of simulations have shown that TSSIM is effective and can assess the jamming effect more accurately in different regions such as complex texture and simple texture area compared with SSIM and PSNR. Shuhong Jiao, Weisheng Dong |
ICIG | 2 |
| 2013 | A learning-based method for compressive image recovery
Weisheng Dong, Guangming Shi, Xiaolin Wu 0001, Lei Zhang 0006 |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | Nonlocal Image Restoration With Bilateral Variance Estimation: A Low-Rank ApproachabstractSimultaneous sparse coding (SSC) or nonlocal image representation has shown great potential in various low-level vision tasks, leading to several state-of-the-art image restoration techniques, including BM3D and LSSC. However, it still lacks a physically plausible explanation about why SSC is a better model than conventional sparse coding for the class of natural images. Meanwhile, the problem of sparsity optimization, especially when tangled with dictionary learning, is computationally difficult to solve. In this paper, we take a low-rank approach toward SSC and provide a conceptually simple interpretation from a bilateral variance estimation perspective, namely that singular-value decomposition of similar packed patches can be viewed as pooling both local and nonlocal information for estimating signal variances. Such perspective inspires us to develop a new class of image restoration algorithms called spatially adaptive iterative singular-value thresholding (SAIST). For noise data, SAIST generalizes the celebrated BayesShrink from local to nonlocal models; for incomplete data, SAIST extends previous deterministic annealing-based solution to sparsity optimization through incorporating the idea of dictionary learning. In addition to conceptual simplicity and computational efficiency, SAIST has achieved highly competent (often better) objective performance compared to several state-of-the-art methods in image denoising and completion experiments. Our subjective quality results compare favorably with those obtained by existing techniques, especially at high noise levels and with a large amount of missing data. Weisheng Dong, Guangming Shi, Xin Li 0005 |
IEEE Trans. Image Process. | 1 |
| 2013 | Sparse Representation Based Image Interpolation With Nonlocal Autoregressive ModelingabstractSparse representation is proven to be a promising approach to image super-resolution, where the low-resolution (LR) image is usually modeled as the down-sampled version of its high-resolution (HR) counterpart after blurring. When the blurring kernel is the Dirac delta function, i.e., the LR image is directly down-sampled from its HR counterpart without blurring, the super-resolution problem becomes an image interpolation problem. In such cases, however, the conventional sparse representation models (SRM) become less effective, because the data fidelity term fails to constrain the image local structures. In natural images, fortunately, many nonlocal similar patches to a given patch could provide nonlocal constraint to the local structure. In this paper, we incorporate the image nonlocal self-similarity into SRM for image interpolation. More specifically, a nonlocal autoregressive model (NARM) is proposed and taken as the data fidelity term in SRM. We show that the NARM-induced sampling matrix is less coherent with the representation dictionary, and consequently makes SRM more effective for image interpolation. Our extensive experimental results demonstrate that the proposed NARM-based image interpolation method can effectively reconstruct the edge structures and suppress the jaggy/ringing artifacts, achieving the best image interpolation results so far in terms of PSNR as well as perceptual quality metrics such as SSIM and FSIM. Weisheng Dong, Lei Zhang 0006, Rastislav Lukac, Guangming Shi |
IEEE Trans. Image Process. | 1 |
| 2013 | Nonlocally Centralized Sparse Representation for Image RestorationabstractSparse representation models code an image patch as a linear combination of a few atoms chosen out from an over-complete dictionary, and they have shown promising results in various image restoration applications. However, due to the degradation of the observed image (e.g., noisy, blurred, and/or down-sampled), the sparse representations by conventional models may not be accurate enough for a faithful reconstruction of the original image. To improve the performance of sparse representation-based image restoration, in this paper the concept of sparse coding noise is introduced, and the goal of image restoration turns to how to suppress the sparse coding noise. To this end, we exploit the image nonlocal self-similarity to obtain good estimates of the sparse coding coefficients of the original image, and then centralize the sparse coding coefficients of the observed image to those estimates. The so-called nonlocally centralized sparse representation (NCSR) model is as simple as the standard sparse representation model, while our extensive experiments on various types of image restoration problems, including denoising, deblurring and super-resolution, validate the generality and state-of-the-art performance of the proposed NCSR algorithm. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xin Li 0005 |
IEEE Trans. Image Process. | 1 |
| 2012 | Image reconstruction with locally adaptive sparsity and nonlocal robust regularization
Weisheng Dong, Guangming Shi, Xin Li 0005, Lei Zhang 0006, Xiaolin Wu 0001 |
Signal Process. Image Commun. | 1 |
| 2012 | Model-Assisted Adaptive Recovery of Compressed Sensing with Imaging ApplicationsabstractIn compressive sensing (CS), a challenge is to find a space in which the signal is sparse and, hence, faithfully recoverable. Since many natural signals such as images have locally varying statistics, the sparse space varies in time/spatial domain. As such, CS recovery should be conducted in locally adaptive signal-dependent spaces to counter the fact that the CS measurements are global and irrespective of signal structures. On the contrary, existing CS reconstruction methods use a fixed set of bases (e.g., wavelets, DCT, and gradient spaces) for the entirety of a signal. To rectify this problem, we propose a new framework for model-guided adaptive recovery of compressive sensing (MARX) and show how a 2-D piecewise autoregressive model can be integrated into the MARX framework to make CS recovery adaptive to spatially varying second order statistics of an image. In addition, MARX offers a mechanism of characterizing and exploiting structured sparsities of natural images, greatly restricting the CS solution space. Simulation results over a wide range of natural images show that the proposed MARX technique can improve the reconstruction quality of existing CS methods by 2-7 dB. Xiaolin Wu 0001, Weisheng Dong, Xiangjun Zhang, Guangming Shi |
IEEE Trans. Image Process. | 2 |
| 2011 | Sparsity-based image denoising via dictionary learning and structural clusteringabstractWhere does the sparsity in image signals come from? Local and nonlocal image models have supplied complementary views toward the regularity in natural images - the former attempts to construct or learn a dictionary of basis functions that promotes the sparsity; while the latter connects the sparsity with the self-similarity of the image source by clustering. In this paper, we present a variational framework for unifying the above two views and propose a new denoising algorithm built upon clustering-based sparse representation (CSR). Inspired by the success of l1-optimization, we have formulated a double-header l1-optimization problem where the regularization involves both dictionary learning and structural structuring. A surrogate-function based iterative shrinkage solution has been developed to solve the double-header l1-optimization problem and a probabilistic interpretation of CSR model is also included. Our experimental results have shown convincing improvements over state-of-the-art denoising technique BM3D on the class of regular texture images. The PSNR performance of CSR denoising is at least comparable and often superior to other competing schemes including BM3D on a collection of 12 generic natural images. Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi |
CVPR | 1 |
| 2011 | Centralized sparse representation for image restorationabstractThis paper proposes a novel sparse representation model called centralized sparse representation (CSR) for image restoration tasks. In order for faithful image reconstruction, it is expected that the sparse coding coefficients of the degraded image should be as close as possible to those of the unknown original image with the given dictionary. However, since the available data are the degraded (noisy, blurred and/or down-sampled) versions of the original image, the sparse coding coefficients are often not accurate enough if only the local sparsity of the image is considered, as in many existing sparse representation models. To make the sparse coding more accurate, a centralized sparsity constraint is introduced by exploiting the nonlocal image statistics. The local sparsity and the nonlocal sparsity constraints are unified into a variational framework for optimization. Extensive experiments on image restoration validated that our CSR model achieves convincing improvement over previous state-of-the-art methods. Weisheng Dong, Lei Zhang 0006, Guangming Shi |
ICCV | 1 |
| 2011 | Sparsity-based image deblurring with locally adaptive and nonlocally robust regularizationabstractImportant structures in photographic images such as edges and textures are jointly characterized by local variation and nonlocal invariance (similarity). Both of them provide valuable heuristics to the regularization of image restoration process. In this pa per, we propose to explore two sets of complementary ideas: 1) locally learn PCA-based dictionaries and estimate the sparsity regularization parameters for each coefficient; and 2) nonlocally enforce the invariance constraint by introducing a patch-similarity based term into the cost functional. The minimization of this new cost functional leads to an iterative thresholding-based image deblurring algorithm and its efficient implementation is discussed. Our experimental results have shown that the proposed scheme significantly outperforms several leading deblurring techniques in the literature on both objective and visual quality assessments. Weisheng Dong, Xin Li 0005, Lei Zhang 0006, Guangming Shi |
ICIP | 1 |
| 2011 | Image Deblurring and Super-Resolution by Adaptive Sparse Domain Selection and Adaptive RegularizationabstractAs a powerful statistical image modeling technique, sparse representation has been successfully used in various image restoration applications. The success of sparse representation owes to the development of the l(1)-norm optimization techniques and the fact that natural images are intrinsically sparse in some domains. The image restoration quality largely depends on whether the employed sparse domain can represent well the underlying image. Considering that the contents can vary significantly across different images or different patches in a single image, we propose to learn various sets of bases from a precollected dataset of example image patches, and then, for a given patch to be processed, one set of bases are adaptively selected to characterize the local sparse domain. We further introduce two adaptive regularization terms into the sparse representation framework. First, a set of autoregressive (AR) models are learned from the dataset of example image patches. The best fitted AR models to a given patch are adaptively selected to regularize the image local structures. Second, the image nonlocal self-similarity is introduced as another regularization term. In addition, the sparsity regularization parameter is adaptively estimated for better image restoration performance. Extensive experiments on image deblurring and super-resolution validate that by using adaptive sparse domain selection and adaptive regularization, the proposed method achieves much better results than many state-of-the-art algorithms in terms of both PSNR and visual perception. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2010 | Live demonstration: Spatial-temporal color video reproduction from noisy CFA sequence track: Digital signal processingabstractThis demonstration shows a spatial-temporal denoising and demosaicking scheme for noisy CFA videos. This scheme can significantly reduce the noise-caused color artifacts and effectively preserve the image edge structures. The experimental results showed that this scheme achieves promising color video reproduction in terms of both PSNR and visual perception. Lei Zhang 0006, Weisheng Dong, Chiu-Wai Hui, Xiaolin Wu 0001, Guangming Shi |
ISCAS | 2 |
| 2010 | Super-resolution with nonlocal regularized sparse representationabstractThe reconstruction of a high resolution (HR) image from its low resolution (LR) counterpart is a challenging problem. The recently developed sparse representation (SR) techniques provide new solutions to this inverse problem by introducing the l1-norm sparsity prior into the super-resolution reconstruction process. In this paper, we present a new SR based image super-resolution by optimizing the objective function under an adaptive sparse domain and with the nonlocal regularization of the HR images. The adaptive sparse domain is estimated by applying principal component analysis to the grouped nonlocal similar image patches. The proposed objective function with nonlocal regularization can be efficiently solved by an iterative shrinkage algorithm. The experiments on natural images show that the proposed method can reconstruct HR images with sharp edges from degraded LR images. Weisheng Dong, Guangming Shi, Lei Zhang 0006, Xiaolin Wu 0001 |
VCIP | 1 |
| 2010 | Two-stage image denoising by principal component analysis with local pixel grouping
Lei Zhang 0006, Weisheng Dong, David Zhang 0001, Guangming Shi |
Pattern Recognit. | 2 |
| 2010 | Spatial-Temporal Color Video Reconstruction From Noisy CFA SequenceabstractSingle-sensor digital video cameras use a color filter array (CFA) to capture video and a color demosaicking (CDM) procedure to reproduce the full color sequence. The reproduced video frames suffer from the inevitable sensor noise introduced in the video acquisition process. This paper presents a spatial-temporal denoising and demosaicking scheme that works without explicit motion estimation. We first perform patch based denoising on the mosaic CFA video. For each CFA patch to be denoised, similar patches are selected within a local spatial-temporal neighborhood. The principal component analysis is performed on the selected patches to remove noise. We then apply an initial single-frame CDM to the denoised CFA data, and subsequently post-process the demosaicked frames by exploiting the spatial-temporal redundancy to reduce the color artifacts. The experimental results on simulated and real noisy CFA sequences demonstrate that the proposed spatial-temporal CFA video denoising and demosaicking scheme can significantly reduce the noise-caused color artifacts and effectively preserve the image edge structures. Lei Zhang 0006, Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Context-based bias removal of statistical models of wavelet coefficients for image denoisingabstractExisting wavelet-based image denoising techniques all assume a probability model of wavelet coefficients that has zero mean, such as zero-mean Laplacian, Gaussian, or generalized Gaussian distributions. While such a zero-mean probability model fits a wavelet subband well, in areas of edges and textures the distribution of wavelet coefficients exhibits a significant bias. We propose a context modeling technique to estimate the expectation of each wavelet coefficient conditioned on the local signal structure. The estimated expectation is then used to shift the probability model of wavelet coefficient back to zero. This bias removal technique can significantly improve the performance of existing wavelet-based image denoisers. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi, Lei Zhang 0006 |
ICIP | 1 |
| 2009 | Nonlocal back-projection for adaptive image enlargementabstractThis paper presents a novel non-local iterative back-projection (NLIBP) algorithm for image enlargement. The iterative back-projection (IBP) technique iteratively reconstructs a high resolution (HR) image from its blurred and downsampled low resolution (LR) counterpart. However, the conventional IBP methods often produce many ¿jaggy¿ and ¿ringing¿ artifacts because the reconstruction errors are back projected into the reconstructed image isotropically and locally. In natural images, usually there exist many non-local redundancies which can be exploited to improve the image reconstruction quality. Therefore, we propose to incorporate adaptively the non-local information into the IBP process so that the reconstruction errors can be reduced. Experimental results demonstrated that the proposed NLBP can reconstruct faithfully the HR images with sharp edges and texture structures. It outperforms the state-of-the-art methods in both PSNR and visual perception. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xiaolin Wu 0001 |
ICIP | 1 |
| 2009 | Learning-based recovery of compressive sensing with application in multiple description codingabstractThe recently proposed compressive sensing (CS) theory provides a new solution for multiple description coding (MDC) with fine granularity, by treating each random CS measurement as a description. The performance of CS-based MDC (CS-MDC) depends on the efficacy of the CS recovery algorithm. Existing CS recovery algorithms recover the signal in a fixed space (e.g., Wavelet, DCT, and gradient spaces) for the entire duration of the signal, even though a typical multimedia signal exhibits sparsity in time/space variant spaces. To rectify this problem and develop a better CS recovery algorithm for CSMDC, we propose a learning-based framework to conduct the CS recovery in locally adaptive spaces, and carry out a case study on image MDC. A set of prior image models are learned offline from a training set to facilitate the CS recovery in local adaptive bases. Experiments show that the learning-based CS recovery algorithm can significantly improve the performance of the previous CS-MDC technique in both PSNR and visual quality. Guangming Shi, Weisheng Dong, Xiaolin Wu 0001 |
MMSP | 3 |
| 2008 | Signal-adapted directional lifting scheme for image compressionabstractIn this paper, we propose a new adaptive lifting scheme that not only locally adapts the filtering directions to the orientations of image features, but also adapts the lifting filters to the statistic properties of image signal. The proposed approach refines previous adaptive directional lifting-based wavelet transform (ADL) by combining directional lifting and adaptive lifting filters to form a unified framework. The image signal is first segmented into regions of textures and edges with close directional features. The lifting filters are then effectively designed. The prediction step is designed to minimize the prediction error of the image signal, and the update step is designed to minimize the reconstruction error. Significant improvements on objective and subjective quality over conventional 2-D wavelet transform and previous ADL transform are achieved. Weisheng Dong, Guangming Shi, Jizheng Xu |
ISCAS | 1 |
| 2008 | Adaptive Nonseparable Interpolation for Image Compression With Directional Wavelet TransformabstractThe adaptive directional lifting-based wavelet transform (ADL) locally adapts the filtering directions to the local properties of the image. In this letter, instead of using the conventional interpolation filter for the directional prediction with fractional-pel accuracy, a new two-dimensional nonseparable adaptive interpolation filter is proposed. The adaptive filter is calculated for every fractional-pel direction so as to minimize the energy of the prediction error. The tradeoff between reducing the prediction error and the overhead to code the interpolation filter is discussed. This enables coding gains of up to 0.98 dB, compared to ADL coder, and up to 2.4 dB, compared to the JPEG 2000 for typical test images. Weisheng Dong, Guangming Shi, Jizheng Xu |
IEEE Signal Process. Lett. | 1 |
| 2007 | Immune memory clonal selection algorithms for designing stack filters
Weisheng Dong, Guangming Shi, Li Zhang 0004 |
Neurocomputing | 1 |