EDBT 2026 Demo / reviewers in the wild / expert
Wenhui Wu 0001
dblp:136/9842-1
· DBLP profile ↗
37ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0002-0416-7719ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Feature Prior-Guided Conditional Diffusion Model for Underwater Image EnhancementabstractUnderwater imaging always suffers from color distortion and reduced visibility due to light absorption and scattering, severely hindering visual perception and analysis. In this letter, we propose an underwater image enhancement framework based on diffusion model augmented with two lightweight guidance modules. The first module is a conditional branch that extracts structural features from a coarsely enhanced version to guide the denoising process toward more faithful restoration. While the second module retrieves high-quality features from a pre-constructed feature dictionary as priors, effectively restoring colors and fine details in degraded regions. Extensive experiments on public underwater image datasets demonstrate that our proposed method outperforms the state-of-the-art approaches both quantitatively and visually. It also generalizes well across various underwater environments, highlighting the effectiveness of incorporating structural and feature-level guidance into the diffusion process. The source code and pre-trained model are available at https://github.com/Juneit/PGUIE. Linwei Zhu, Tao Tian, Wenhui Wu 0001, Jingchao Cao |
IEEE Signal Process. Lett. | 4 |
| 2025 | GeCC: Generalized Contrastive Clustering with Domain Shifts ModelingabstractContrastive clustering performs clustering and data representation in a unified model, where instance- and cluster-level constrastive learning are conducted simultaneously. However, commonly-used data augmentation methods make contrastive mechanism effect but may cause representation learning getting stuck in domain-specific information, which further deteriorates clustering performance and limits generalization ability. To this end, we propose a new framework, named Generalized Contrastive Clustering with domain shifts modeling (GeCC), which can integrate diverse domain knowledge to improve the clustering performance. Specifically, we first design a cluster-guided domain shifts modeling module to synthesize a reference view with diverse domain information. Then, we introduce instance representation and cluster assignment contrastive modules with well-designed attention weights to guide the representation learning and clustering. In this way, our method can maximize the extraction of cluster-related information and avoid over-fitting domain-specific features. Experimental results on four benchmark datasets demonstrate that our proposed method consistently outperforms other state-of-the-art methods. Wenhui Wu 0001, Le Ou-Yang, Ran Wang 0001, Debby Dan Wang |
AAAI | 2 |
| 2025 | OptiDiff: Unsupervised Deep-Sea Image Enhancement via Optical Priors Guided Stable DiffusionabstractDeep-sea images suffer from extreme light attenuation and non-uniform illumination caused by artificial light sources. To get rid of limitation aroused by low-quality training data, we propose an unsupervised approach for deep-sea image enhancement based on the prior-to-image framework, termed as OptiDiff. Instead of learning mapping from paired underwater dataset, the framework for OptiDiff is trained on air image dataset. Specifically, four optical-invariant priors (OIPs) are used for guiding the stable diffusion model to recover degraded underwater image. One of the utilized OIPs is particularly designed for recover blurred details in background. Besides, to simulate the domain shift between air and underwater images, a channel attenuation strategy derived from characteristics of real-world underwater image is equipped with the framework. Extensive experimental results demonstrate the superiority of the proposed OptiDiff on restoring image from low contrast, low visibility, and severe blur, also showing robustness on shallow-water datasets. Code is available at https://github.com/Miaaaaaa1024/OptiDiff. Wenhui Wu 0001, Yuemiao Wang, Hua Li 0012, Yuanhao Gong |
ICME | 1 |
| 2025 | DUIMC: Deep Unbalanced Incomplete Multi-View Clustering via Graph Constrained Imputation and Contrastive LearningabstractDue to the frequent occurrence of missing views in real-world multi-view data, incomplete multi-view clustering (IMVC) has attracted significant attention. However, most existing IMVC methods overlook the fact that incomplete data in practical applications often exhibits varying missing rates across different views, rendering their mechanisms ineffective under such conditions. Although several works based on conventional learning methods have been proposed to solve unbalanced incomplete multi-view clustering (UIMVC), their performance is limited by their shallow feature representation and over-sophisticated optimization procedure. In this paper, we propose Deep Unbalanced Incomplete Multi-view Clustering via Graph Constrained Imputation and Contrastive Learning (DUIMC) to address UIMVC with deep learning paradigm. Specifically, DUIMC introduces a novel differentiable imputation layer for dynamically handling unbalanced incompleteness and integrates it with multi-view contrastive clustering into a unified deep representation learning framework. Furthermore, bi-level graph constraints are imposed on imputation and representation learning to preserve local consistency at both the feature and instance levels. In addition, we develop adaptive fusion mechanisms to adaptively restrain the impact aroused by information unbalance among views. Extensive experimental results on five benchmark datasets demonstrate DUIMC's superior clustering performance over several traditional state-of-the-art approaches. Wenhui Wu 0001, Guanqi Wen, Le Ou-Yang, Ran Wang 0001, Sam Kwong |
ACM Multimedia | 1 |
| 2025 | Interpretable Optimization-Inspired Unfolding Network for Low-Light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement (LLIE). However, the hand-crafted priors and conventional optimization algorithm adopted to solve the layer decomposition problem result in the lack of adaptivity and efficiency. To this end, this paper proposes a Retinex-based deep unfolding network (URetinex-Net++), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and fairly-flexible component adjustment, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in the data-driven manner, can realize noise suppression and details preservation for decomposed components. URetinex-Net++ is a further augmented version of URetinex-Net, which introduces a cross-stage fusion block to alleviate the color defect in URetinex-Net. Therefore, boosted performance on LLIE can be obtained in both visual quality and quantitative metrics, where only a few parameters are introduced and little time is cost. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed URetinex-Net++ over state-of-the-art methods. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Semi-Supervised Symmetric Non-Negative Matrix Factorization With Low-Rank Tensor RepresentationabstractSemi-supervised symmetric non-negative matrix factorization (SNMF) utilizes the available supervisory information (usually in the form of pairwise constraints) to improve the clustering ability of SNMF. The previous methods introduce the pairwise constraints from the local perspective, i.e., they either directly refine the similarity matrix element-wisely or restrain the distance of the decomposed vectors in pairs according to the pairwise constraints, which overlook the global perspective, i.e., in the ideal case, the pairwise constraint matrix and the ideal similarity matrix possess the same low-rank structure. To this end, we first propose a novel semi-supervised SNMF model by seeking low-rank representation for the tensor synthesized by the pairwise constraint matrix and a similarity matrix obtained by the product of the embedding matrix and its transpose, which could strengthen those two matrices simultaneously from a global perspective. We then propose an enhanced SNMF model, making the embedding matrix tailored to the above tensor low-rank representation. We finally refine the similarity matrix by the strengthened pairwise constraints. We repeat the above steps to continuously boost the similarity matrix and pairwise constraint matrix, leading to a high-quality embedding matrix. Extensive experiments substantiate the superiority of our method. The code is available athttps://github.com/JinaLeejnl/TSNMF. Yuheng Jia, Wenhui Wu 0001, Ran Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Low-Light Image Enhancement Through Learning a Simplified Inverse Rendering ModelabstractIt remains to be extremely difficult to capture high quality photographs of low-light scenes. Low light causes the low signal-to-noise ratio (SNR) problem which makes the image noisy. Such scenes almost always have the high dynamic range (HDR) problem caused by uneven lighting where a small area surrounding the light source is very bright while the rest of the scene is very dark, making it very difficult to simultaneously obtain high quality signals in both the dark and bright areas. This paper presents a new image restoration method for tackling the problems in low-light scenes. Fundamentally differing from existing approaches, the new method borrows ideas from inverse graphics rendering and re-renders the image with a canonical light source thus correcting the image from first principle. A deep learning based simplified inverse rendering model (SIRM) featuring implicit regularization is first developed for correcting uneven lighting and then an end-to-end convolutional neural network is constructed for reducing noise. Extensive experimental results are presented to demonstrate that the new method outperforms state-of-the-art methods, and is capable of effectively brightening up dark image regions while at the same time preserving details and color consistency. Our code is available at: https://github.com/pj0927/SIRNet. Wenhui Wu 0001, Jia Pang, Shuaibo Gao, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | GRESS: Grouping Belief-Based Deep Contrastive Subspace ClusteringabstractThe self-expressive coefficient plays a crucial role in the self-expressiveness-based subspace clustering method. To enhance the precision of the self-expressive coefficient, we propose a novel deep subspace clustering method, named grouping belief-based deep contrastive subspace clustering (GRESS), which integrates the clustering information and higher-order relationship into the coefficient matrix. Specifically, we develop a deep contrastive subspace clustering module to enhance the learning of both self-expressive coefficients and cluster representations simultaneously. This approach enables the derivation of relatively noiseless self-expressive similarities and cluster-based similarities. To enable interaction between these two types of similarities, we propose a unique grouping belief-based affinity refinement module. This module leverages grouping belief to uncover the higher-order relationships within the similarity matrix, and integrates the well-designed noisy similarity suppression and similarity increment regularization to eliminate redundant connections while complete absent information. Extensive experimental results on four benchmark datasets validate the superiority of our proposed method GRESS over several state-of-the-art methods. Wenhui Wu 0001, Le Ou-Yang, Ran Wang 0001, Sam Kwong |
IEEE Trans. Cybern. | 2 |
| 2025 | Weakly-Supervised 3D Visual Grounding Based on Visual Language AlignmentabstractLearning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial number of bounding box annotations for text queries, which is time-consuming and labor-intensive to obtain. In this paper, we propose3D-VLA, a weakly supervised approach for3Dvisual grounding based onVisualLanguageAlignment. Our 3D-VLA exploits the superior ability of current large-scale vision-language models (VLMs) on aligning the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds with no need for fine-grained box annotations in the training procedure. During the inference stage, the learned text-3D correspondence will help us ground the text queries to the 3D target objects even without 2D images. To the best of our knowledge, this is the first work to investigate 3D visual grounding in a weakly supervised manner by involving large scale vision-language models, and extensive experiments on ReferIt3D and ScanRefer datasets demonstrate that our 3D-VLA achieves comparable and even superior results over the fully supervised methods. Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu 0001, Zequn Jie, Lin Ma 0002, Xu Wang 0006 |
IEEE Trans. Multim. | 4 |
| 2025 | HNR-ISC: Hybrid Neural Representation for Image Set CompressionabstractImage set compression (ISC) refers to compressing the sets of semantically similar images. Traditional ISC methods typically aim to eliminate redundancy among images at either signal or frequency domain, but often struggle to handle complex geometric deformations across different images effectively. Here, we propose a new Hybrid Neural Representation for ISC (HNR-ISC), including an implicit neural representation for Semantically Common content Compression (SCC) and an explicit neural representation for Semantically Unique content Compression (SUC). Specifically, SCC enables the conversion of semantically common contents into a small-and-sweet neural representation, along with embeddings that can be conveyed as a bitstream. SUC is composed of invertible modules for removing intra-image redundancies. The feature level combination from SCC and SUC naturally forms the final image set. Experimental results demonstrate the robustness and generalization capability of HNR-ISC in terms of signal and perceptual quality for reconstruction and accuracy for the downstream analysis task. Shiqi Wang 0001, Meng Wang 0017, Peilin Chen 0001, Wenhui Wu 0001, Xu Wang 0006, Sam Kwong |
IEEE Trans. Multim. | 5 |
| 2025 | RGB-D Data Compression via Bi-Directional Cross-Modal Prior Transfer and Enhanced Entropy ModelingabstractRGB-D data, being homogeneous cross-modal data, demonstrates significant correlations among data elements. However, current research focuses only on a uni-directional pattern of cross-modal contextual information, neglecting the exploration of bi-directional relationships in the compression field. Thus, we propose a joint RGB-D compression scheme, which is combined with Bi-Directional Cross-Modal Prior Transfer (Bi-CPT) modules and a Bi-Directional Cross-Modal Enhanced Entropy (Bi-CEE) model. The Bi-CPT module is designed for compact representations of cross-modal features, effectively eliminating spatial and modality redundancies at different granularity levels. In contrast to the traditional entropy models, our proposed Bi-CEE model not only achieves spatial-channel contextual adaptation through partitioning RGB and depth features but also incorporates information from other modalities as prior to enhance the accuracy of probability estimation for latent variables. Furthermore, this model enables parallel multi-stage processing to accelerate coding. Experimental results demonstrate the superiority of our proposed framework over the current compression scheme, outperforming both rate-distortion performance and downstream tasks, including surface reconstruction and semantic segmentation. The source code will be available at https://github.com/xyy7/Learning-based-RGB-D-Image-Compression . Yuyu Xu, Qiudan Zhang, Wenhui Wu 0001, Yun Zhang 0002, Xu Wang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | On the Adversarial Robustness of Hierarchical ClassificationabstractDeep neural networks (DNNs) have demonstrated remarkable success on various learning problems, but they face a formidable challenge in the form of adversarial attacks. Especially, when dealing with complex classification tasks for numerous classes with a hierarchical structure, the adversarial robustness of a DNN model may drop seriously. In this paper, we investigate the adversarial robustness of DNN models on such complex classification tasks. In response, we propose a two-stage hierarchical classification framework, which is composed of a coarse-grained classifier and a series of fine-grained classifiers. A data correction sampling module is designed between the two stages, in order to mitigate the influence of misclassification caused by the coarse-grained classifier; and a discriminative filter learning module is employed in the fine-grained classification, in order to gain better distinguish abilities among fine-grained categories. Experiments on the well-known dataset CIFAR-100 and a newly-constructed hierarchical dataset mini-ImageNet76 demonstrate that employing a hierarchical framework can effectively improve the model robustness on such complex classification tasks. Ran Wang 0001, Simeng Zeng, Wenhui Wu 0001, Yuheng Jia, Wing W. Y. Ng, Xizhao Wang |
SMC | 3 |
| 2024 | Adversarially robust neural networks with feature uncertainty learning and label embedding
Ran Wang 0001, Haopeng Ke, Wenhui Wu 0001 |
Neural Networks | 4 |
| 2024 | Image Intrinsic Components Guided Conditional Diffusion Model for Low-Light Image EnhancementabstractThrough formulating the image restoration as a generation problem, the conditional diffusion model has been applied to low-light image enhancement (LIE) to restore the details in dark regions. However, in the previous diffusion model based LIE methods, the conditions used for guiding generation are degraded images, such as low-light image, signal-to-noise ratio map and color map, which suffer from severe degradation and are simply fed into diffusion model by rigidly concatenating with the noise. To avoid using degraded conditions resulting in sub-optimal performance in recovering details and enhancing brightness, we use the image intrinsic components originating from the Retinex model as guidance, whose multi-scale features are flexibly integrated into the diffusion model, and propose a novel conditional diffusion model for LIE. Specifically, the input low-light image is decomposed into reflectance and illumination by a Retinex decomposition module, where two components contain abundant physical property and lighting conditions of the scene. Then, we extract the latent features from two conditions through a component-dependent feature extraction module, which is designed according to the physical property of components. Finally, instead of previous rigid concatenation manner, a well-designed feature fusion mechanism is equipped to adaptively embed generative conditions into diffusion model. Extensive experimental results demonstrate that our method outperforms the state-of-the-art methods, and is capable of effectively restoring the local details while brightening the dark regions. Our codes are available athttps://github.com/Knossosc/ICCDiff. Sicong Kang, Shuaibo Gao, Wenhui Wu 0001, Xu Wang 0006, Shuoyao Wang, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Neural Network Based Multi-Level In-Loop Filtering for Versatile Video CodingabstractTo further improve the performance of Versatile Video Coding (VVC), a neural network based multi-level in-loop filtering framework for luma and chroma is presented in this letter, which includes Reference pixel Level (RL), Coding tree unit Level (CL), and Frame Level (FL). The neural network based filters in these levels can be flexibly enabled. In RL, the coding performance upper bound is analyzed and asymmetric convolution is designed. In CL, the pixels located at the bottom and rightmost have been assigned greater weights for loss calculation during training. In addition, the co-located luma is adopted in CL and FL chroma filtering for guiding chroma enhancement due to the high correlation between them. For the architecture of neural network, two input channel fusion schemes are combined to enjoy both of their benefits. Extensive experimental results show that the proposed multi-level in-loop filtering method can achieve 6.87%, 32.8%, and 36.9% bit rate reductions on average for Y, U, and V components under all intra configuration, which outperforms the state-of-the-art works. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Wenhui Wu 0001, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Stereo Superpixel Segmentation via Decoupled Dynamic Spatial-Embedding Fusion NetworkabstractStereo superpixel segmentation aims at grouping the discretizing pixels into perceptual regions through left and right views more collaboratively and efficiently. Existing superpixel segmentation algorithms mostly utilize color and spatial features as input, which may impose strong constraints on spatial information while utilizing the disparity information in terms of stereo image pairs. To alleviate this issue, we propose a stereo superpixel segmentation method with a decoupling mechanism of spatial information in this work. To decouple stereo disparity information and spatial information, the spatial information is temporarily removed before fusing the features of stereo image pairs, and a decoupled stereo fusion module (DSFM) is designed to handle the stereo features alignment as well as occlusion problems. Moreover, since the spatial information is vital to superpixel segmentation, we further design a dynamic spatiality embedding module (DSEM) to re-add spatial information, and the weights of spatial information will be adaptively adjusted through the dynamic fusion (DF) mechanism in DSEM for achieving a finer segmentation. Comprehensive experimental results demonstrate that our method can achieve the state-of-the-art performance on the KITTI2015 and Cityscapes datasets, and also verify the efficiency when applied in salient object detection on NJU2K dataset. The source code will be available publicly after paper is accepted. Hua Li 0012, Junyan Liang, Runmin Cong, Wenhui Wu 0001, Sam Kwong |
IEEE Trans. Multim. | 5 |
| 2024 | Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-LabelingabstractLearning to build 3D scene graphs is essential for real-world perception in a structured and rich fashion. However, previous 3D scene graph generation methods utilize a fully supervised learning manner and require a large amount of entity-level annotation data of objects and relations, which is extremely resource-consuming and tedious to obtain. To tackle this problem, we propose 3D-VLAP, a weakly-supervised 3D scene graph generation method via Visual-Linguistic Assisted Pseudo-labeling. Specifically, our 3D-VLAP exploits the superior ability of current large-scale visual-linguistic models to align the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds. First, we establish the positional correspondence from 3D point clouds to 2D images via camera intrinsic and extrinsic parameters, thereby achieving alignment of 3D point clouds and 2D images. Subsequently, a large-scale cross-modal visual-linguistic model is employed to indirectly align 3D instances with the textual category labels of objects by matching 2D images with object category labels. The pseudo labels for objects and relations are then produced for 3D-VLAP model training by calculating the similarity between visual embeddings and textual category embeddings of objects and relations encoded by the visual-linguistic model, respectively. Ultimately, we design an edge self-attention based graph neural network to generate scene graphs of 3D point clouds. Experiments demonstrate that our 3D-VLAP achieves comparable results with current fully supervised methods, meanwhile alleviating the data annotation pressure. Xu Wang 0006, Qiudan Zhang, Wenhui Wu 0001, Mark Junjie Li, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 4 |
| 2023 | Hybrid Prior-Based Diminished Reality for Indoor Panoramic Images
Jiashu Liu, Qiudan Zhang, Xuelin Shen, Wenhui Wu 0001, Xu Wang 0006 |
CGI (3) | 4 |
| 2023 | A Multi-Stream Network for Mesh Denoising Via Graph Neural Networks with Gaussian Curvatureabstract3D meshes are getting popular in both research and industry. However, the meshes obtained via the 3D scanning equipment frequently contain a high level of noise. In this paper, we present a Gaussian Curvature Driven Multi-stream Network (GCM-Net) based on graph convolutional networks. This network can remove the noise while preserving the essential features during the 3D mesh denoising process. Our method is the first attempt to apply the high-order feature (i.e., Gaussian curvature) in the denoising task, which is more descriptive for the shape of the mesh. GCM-Net consists of curvature stream, vertex stream, and face normal stream, where the curvature stream focuses on the high-order Gaussian curvature feature of 3D mesh. Our method achieves state-of-the-art results on a publicly available dataset, demonstrating its effectiveness. The proposed method can be applied in various applications, such as 3D human body modeling, metaverse, object tracking and biomedical visualization. Zhibo Zhao, Wenhui Wu 0001, Yuanhao Gong |
ICIP | 2 |
| 2023 | Multi-View Super Resolution for Underwater Images Utilizing Atmospheric Light Scattering ModelabstractThe underwater environment is complex and the underwater light propagation undergoes absorption, scattering and reflection. This leads to the fact that the underwater light imaging cannot be generalized from land-based. How to use these imaging features to work better with super-resolution tasks for underwater imagery applications is still rarely studied. In this paper, we introduce the medium transmission (MT) maps to advance super-resolution tasks for underwater images. A multi-view network is designed to fuse information from the original underwater images and the MT maps, which provides information on the underlying physical properties of the water, such as the attenuation coefficients in different parts of water. By integrating information from multiple views, the proposed network can capture more of the underlying structure and features of the scene, leading to higher-quality super-resolved images. Besides, a new loss function, namely MT Loss, is developed according to the lack of details in special region of the underwater images. This loss function emphasizes the regions with less influence from the underwater environment during the underwater imaging process and therefore the network outputs a more detailed image. Finally, we compare our algorithm with state-of-the-art methods, and extensive results show that our network achieves better qualitative and quantitative performance. Jin Hao, Wenli Duan, Guangfei Li, Shiyan Chen, Wenhui Wu 0001, Hua Li 0012 |
ICPADS | 5 |
| 2023 | FSNet: Frequency Domain Guided Superpixel Segmentation Network for Complex ScenesabstractExisting superpixel segmentation algorithms mainly focus on natural image with high-quality, while neglecting the inevitable environment constraint in complex scenes. In this paper, we propose an end-to-end frequency domain guided superpixel segmentation network (FSNet) to generate superpixels with sharp boundary adherence for complex scenes by fusing the deep features in spatial and frequency domains. To utilize the frequency domain information of the image, an improved frequency information extractor (IFIE) is proposed to extract the frequency domain information with sharp boundary features. Moreover, considering the over-sharp feature may damage the semantic information of superpixel, we further design a dense hybrid atrous convolution (DHAC) block to preserve semantic information via capturing wider and deeper semantic information in spatial domain. Finally, the extracted deep features in spatial and frequency domains will be fused to generate semantic perceptual superpixels with sharp boundary adherence. Extensive experiments on multiple challenging datasets with complex boundaries demonstrate that our method achieves the state-of-the-art performance both quantitatively and qualitatively, and we further verify the superiority of the proposed method when applied in salient object detection. Hua Li 0012, Junyan Liang, Wenhui Wu 0001 |
ACM Multimedia | 4 |
| 2023 | Improving Federated Person Re-Identification through Feature-Aware Proximity and AggregationabstractPerson re-identification (ReID) is a challenging task that aims to identify individuals across multiple non-overlapping camera views. To enhance the performance and robustness of ReID models, it is crucial to train them over multiple data sources. However, the traditional centralized approach poses a significant challenge to privacy as it requires collecting data from distributed data owners. To overcome this challenge, we employ the federated learning approach, which enables distributed model training without compromising data privacy. In this paper, we propose a novel feature-aware local proximity and global aggregation method for federated ReID to extract robust feature representations. Specifically, we introduce a proximal term and a feature regularization term for local model training to improve local training accuracy while ensuring global aggregation convergence. Furthermore, we use the cosine distance of backbone features to determine the global aggregation weight of each local model. Our proposed method significantly improves the performance and generalization of the global model. Extensive experiments demonstrate the effectiveness of our proposal. Specifically, our method achieves an additional 27.3% Rank-1 average accuracy in federated full supervision and an extra 20.3% mean Average Precision (mAP) on DukeMTMC in federated domain generalization. Pengling Zhang, Huibin Yan, Wenhui Wu 0001, Shuoyao Wang |
ACM Multimedia | 3 |
| 2023 | Self-representative kernel concept factorization
Wenhui Wu 0001, Ran Wang 0001, Le Ou-Yang |
Knowl. Based Syst. | 1 |
| 2023 | Semi-supervised adaptive kernel concept factorization
Wenhui Wu 0001, Junhui Hou, Shiqi Wang 0001, Sam Kwong, Yu Zhou 0027 |
Pattern Recognit. | 1 |
| 2023 | Atmospheric Scattering Model Induced Statistical Characteristics Estimation for Underwater Image RestorationabstractUnderwater images often suffer from color deviation and low contrast due to selective absorption and light scattering, whose degradation is generally described by an Atmospheric Scattering Model (ASM). However, it is challenging to design hand-craft priors to estimate the transmission map and global light within ASM. To avoid the estimation on these two variables, in this paper, we establish a statistical characteristics relationship between underwater and recovered images based on ASM. With this relationship, a novel lightweight model is proposed for efficient Underwater Image Restoration (UIR). Within our proposed model, the UIR problem is disentangled into global restoration and local compensation, for which two modules are developed. Extensive experimental results demonstrate that our proposed method can effectively improve color deviation and low contrast while preserving details, and outperform state-of-the-art methods. Shuaibo Gao, Wenhui Wu 0001, Hua Li 0012, Linwei Zhu, Xu Wang 0006 |
IEEE Signal Process. Lett. | 2 |
| 2022 | URetinex-Net: Retinex-based Deep Unfolding Network for Low-light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement. However, the commonly used handcrafted priors and optimization-driven solutions lead to the absence of adaptivity and efficiency. To address these issues, in this paper, we propose a Retinex-based deep unfolding network (URetinex-Net), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and user-specified illumination enhancement, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in data-driven manner, can realize noise suppression and details preservation for the final decomposition results. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed method over state-of-the-art methods. The code is available at https://github.com/AndersonYong/URetinex-Net. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
CVPR | 1 |
| 2022 | Self-supervised Indoor 360-Degree Depth Estimation via Structural Regularization
Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Wenhui Wu 0001, Xu Wang 0006 |
PRICAI (3) | 5 |
| 2021 | Structure Adaptive Filtering for Edge-Preserving Image Smoothing
Wenming Tang, Yuanhao Gong, Linyu Su, Wenhui Wu 0001, Guoping Qiu |
ICIG (3) | 4 |
| 2021 | Positive and Negative Label-Driven Nonnegative Matrix FactorizationabstractPositive label is often used as the supervisory information in the learning scenario, which refers to the category that a sample is assigned to. However, another side information lying in the labels, which describes the categories that a sample is exclusive of, have been largely ignored. In this paper, we propose a nonnegative matrix factorization (NMF) based classification method leveraging both positive and negative label information, which is termed as positive and negative label-driven NMF (PNLD-NMF). The proposed scheme concurrently accomplishes data representation and classification in a joint manner. Owing to the complementary characteristics between positive and negative labels, we further design a new regularization framework to take advantage of these two label types. Extensive experiments on six image classification benchmark datasets show that the proposed scheme is able to consistently deliver better classification accuracy. Wenhui Wu 0001, Yuheng Jia, Shiqi Wang 0001, Ran Wang 0001, Hongfei Fan, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Superpixel Segmentation Based on Spatially Constrained Subspace ClusteringabstractSuperpixel segmentation aims at dividing the input image into some representative regions containing pixels with similar and consistent intrinsic properties, without any prior knowledge about the shape and size of each superpixel. In this article, to alleviate the limitation of superpixel segmentation applied in practical industrial tasks that detailed boundaries are difficult to be kept, we regard each representative region with independent semantic information as a subspace, and correspondingly formulate superpixel segmentation as a subspace clustering problem to preserve more detailed content boundaries. We show that a simple integration of superpixel segmentation with the conventional subspace clustering does not effectively work due to the spatial correlation of the pixels within a superpixel, which may lead to boundary confusion and segmentation error when the correlation is ignored. Consequently, we devise a spatial regularization and propose a novel convex locality-constrained subspace clustering model that is able to constrain the spatial adjacent pixels with similar attributes to be clustered into a superpixel and generate the content-aware superpixels with more detailed boundaries. Finally, the proposed model is solved by an efficient alternating direction method of multipliers solver. Experiments on different standard datasets demonstrate that the proposed method achieves superior performance both quantitatively and qualitatively compared with some state-of-the-art methods. Hua Li 0012, Yuheng Jia, Runmin Cong, Wenhui Wu 0001, Sam Kwong, Chuanbo Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Joint Optimization for Pairwise Constraint PropagationabstractConstrained spectral clustering (SC) based on pairwise constraint propagation has attracted much attention due to the good performance. All the existing methods could be generally cast as the following two steps, i.e., a small number of pairwise constraints are first propagated to the whole data under the guidance of a predefined affinity matrix, and the affinity matrix is then refined in accordance with the resulting propagation and finally adopted for SC. Such a stepwise manner, however, overlooks the fact that the two steps indeed depend on each other, i.e., the two steps form a "chicken-egg" problem, leading to suboptimal performance. To this end, we propose a joint PCP model for constrained SC by simultaneously learning a propagation matrix and an affinity matrix. Especially, it is formulated as a bounded symmetric graph regularized low-rank matrix completion problem. We also show that the optimized affinity matrix by our model exhibits an ideal appearance under some conditions. Extensive experimental results in terms of constrained SC, semisupervised classification, and propagation behavior validate the superior performance of our model compared with state-of-the-art methods. Yuheng Jia, Wenhui Wu 0001, Ran Wang 0001, Junhui Hou, Sam Kwong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Semi-Supervised Non-Negative Matrix Factorization With Dissimilarity and Similarity RegularizationabstractIn this article, we propose a semi-supervised non-negative matrix factorization (NMF) model by means of elegantly modeling the label information. The proposed model is capable of generating discriminable low-dimensional representations to improve clustering performance. Specifically, a pair of complementary regularizers, i.e., similarity and dissimilarity regularizers, is incorporated into the conventional NMF to guide the factorization. And, they impose restrictions on both the similarity and dissimilarity of the low-dimensional representations of data samples with labels as well as a small number of unlabeled ones. The proposed model is formulated as a well-posed constrained optimization problem and further solved with an efficient alternating iterative algorithm. Moreover, we theoretically prove that the proposed algorithm can converge to a limiting point that meets the Karush-Kuhn-Tucker conditions. Extensive experiments as well as comprehensive analysis demonstrate that the proposed model outperforms the state-of-the-art NMF methods to a large extent over five benchmark data sets, i.e., the clustering accuracy increases to 82.2% from 57.0%. Yuheng Jia, Sam Kwong, Junhui Hou, Wenhui Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Sparse Bayesian Learning-Based Kernel Poisson RegressionabstractIn this paper, we introduce a closed-form sparse Bayesian kernel Poisson regression (SBKPR) model for count data regression problems based on the sparse Bayesian learning (SBL) approach. In Bayesian setting, a Gaussian prior is given to the model parameter, which is not the conjugate distribution of Poisson regression. Hence, the model parameters cannot be integrated analytically, which leads to the inference intractable problem. In this paper, the log-gamma Gaussian approximation method is proposed to solve this analytically intractable problem, which can give out the closed-form solutions. Furthermore, an individual Gaussian prior is given to the model parameters, which can enhance the flexibility of the proposed method. Finally, sparse solutions can be obtained by applying SBL, which can benefit the learning efficiency and reduce the computational time in practical applications. Experimental results demonstrate that the proposed SBKPR model can outperform some state-of-the-art count data regression models on both toy data and real-world data. Yuheng Jia, Sam Kwong, Wenhui Wu 0001, Ran Wang 0001, Wei Gao 0003 |
IEEE Trans. Cybern. | 3 |
| 2019 | Simultaneous Dimensionality Reduction and Classification via Dual Embedding Regularized Nonnegative Matrix FactorizationabstractNonnegative matrix factorization (NMF) is a well-known paradigm for data representation. Traditional NMF-based classification methods first perform NMF or one of its variants on input data samples to obtain their low-dimensional representations, which are successively classified by means of a typical classifier [e.g., k -nearest neighbors (KNN) and support vector machine (SVM)]. Such a stepwise manner may overlook the dependency between the two processes, resulting in the compromise of the classification accuracy. In this paper, we elegantly unify the two processes by formulating a novel constrained optimization model, namely dual embedding regularized NMF (DENMF), which is semi-supervised. Our DENMF solution simultaneously finds the low-dimensional representations and assignment matrix via joint optimization for better classification. Specifically, input data samples are projected onto a couple of low-dimensional spaces (i.e., feature and label spaces), and locally linear embedding is employed to preserve the identical local geometric structure in different spaces. Moreover, we propose an alternating iteration algorithm to solve the resulting DENMF, whose convergence is theoretically proven. Experimental results over five benchmark datasets demonstrate that DENMF can achieve higher classification accuracy than state-of-the-art algorithms. Wenhui Wu 0001, Sam Kwong, Junhui Hou, Yuheng Jia, Horace Ho-Shing Ip |
IEEE Trans. Image Process. | 1 |
| 2018 | Convex Constrained Clustering with Graph-Laplacian PcaabstractIn this paper, we propose a new algorithm for constrained clustering, in which a new regularizer elegantly incorporates a small amount of weakly supervisory information in the form of pair-wise constraints to regularize the similarity between the low-dimensional representations of a set of data samples. By exploring both the local and global structures of the data samples with the guidance of the supervisory information, the proposed algorithm is capable of learning the low-dimensional representations with strong separability. Technically, the proposed algorithm is formulated and relaxed as a convex optimization model, which is further efficiently solved with the global convergence guaranteed. Experimental results on multiple benchmark data sets show that our proposed model can produce higher clustering accuracy than state-of-the-art algorithms. Yuheng Jia, Sam Kwong, Junhui Hou, Wenhui Wu 0001 |
ICME | 4 |
| 2018 | Nonnegative matrix factorization with mixed hypergraph regularization for community detection
Wenhui Wu 0001, Sam Kwong, Yu Zhou 0027, Yuheng Jia, Wei Gao 0003 |
Inf. Sci. | 1 |
| 2018 | Pairwise Constraint Propagation-Induced Symmetric Nonnegative Matrix FactorizationabstractAs a variant of nonnegative matrix factorization (NMF), symmetric NMF (SNMF) has shown to be effective for capturing the cluster structure embedded in the graph representation. In contrast to the existing SNMF-based clustering methods that empirically construct the similarity matrix and rigidly introduce the supervisory information to the assignment matrix, in this paper, we propose a novel SNMF-based semisupervised clustering method, namely, pairwise constraint propagation-induced SNMF (PCPSNMF). By formulating a single-constrained optimization problem, PCPSNMF is capable of learning the similarity and assignment matrices adaptively and simultaneously, in which a small amount of supervisory information in the form of pairwise constraints is introduced in a flexible way to guide the construction of the similarity matrix, and the two matrices communicate with each other to achieve mutual refinement until convergence. In addition, we propose an efficient alternating iterative algorithm to solve the optimization problem, whose convergence is theoretically proven. Experimental results over several benchmark image data sets demonstrate that PCPSNMF is less sensitive to initialization and produces higher clustering performance, compared with the state-of-the-art methods. Wenhui Wu 0001, Yuheng Jia, Sam Kwong, Junhui Hou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |