Cong Wang 0018

dblp:18/2771-18 · DBLP profile ↗
← Back
48ranked-venue papers
20as first author
39since 2021 · last 2026
0000-0002-6068-0103ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 15 first-author · 23 since 2021Artificial intelligence and machine learning · 27 · 10 first-author · 24 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
abstract
Image-based virtual try-on (VTON) aims to synthesize photorealistic images of a person wearing specified garments. Despite significant progress, building a universal VTON framework that can flexibly handle diverse and complex tasks remains a major challenge. Recent methods explore multi-task VTON frameworks guided by textual instructions, yet they still face two key limitations: (1) semantic gap between text instructions and reference images, and (2) data scarcity in complex scenarios. To address these challenges, we propose UniFit, a universal VTON framework driven by a Multimodal Large Language Model (MLLM). Specifically, we introduce an MLLM-Guided Semantic Alignment Module (MGSA), which integrates multimodal inputs using an MLLM and a set of learnable queries. By imposing a semantic alignment loss, MGSA captures cross-modal semantic relationships and provides coherent and explicit semantic guidance for the generative process, thereby reducing the semantic gap. Moreover, by devising a two-stage progressive training strategy with a self-synthesis pipeline, UniFit is able to learn complex tasks from limited data. Extensive experiments show that UniFit not only supports a wide range of VTON tasks, including multi-garment and model-to-model try-on, but also achieves state-of-the-art performance.
Wei Zhang 0196, Yeying Jin, Xin Li 0082, Yan Zhang 0004, Xiaofeng Cong, Cong Wang 0018, Fengcai Qiao, Zhichao Lian
AAAI6
2026 Neural Discrimination-Prompted Transformers for Efficient UHD Image Restoration and Enhancement
Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Yang Yang 0009
Int. J. Comput. Vis.1
2026 Phrase Grounding-Based Style Transfer for Single-Domain Generalized Object Detection
abstract
Single-domain generalized object detection aims to enhance a model’s generalization to multiple unseen target domains using only data from a single source domain during training. This is a practical yet challenging scenario, as it requires the model to address domain shift without incorporating target domain data into the training process. In this paper, we propose a novel phrase-grounding-based style transfer (PGST) approach for the task. Specifically, we first define textual prompts to describe objects for potential unseen target domains. Then, we leverage the grounded language-image pre-training (GLIP) model to capture the styles of these target domains and perform style transfer from the source to the target domains. The style-transferred visual features from the source domain are semantically rich and closely approximate those of their hypothetical counterparts in the target domain. Finally, we employ these style-transferred visual features to fine-tune GLIP. By introducing these imaginary counterparts, the detector can be effectively generalized to unseen target domains using only a single source domain during training. Our method significantly improves mean average precision (mAP), with an average increase of 8.8% across five diverse weather-driving benchmarks. Notably, our approach outperforms or matches the performance of domain-adaptive object detection methods, which require target domain data for training, in several challenging scenarios.
Wei Wang 0335, Cong Wang 0018, Mengzhu Wang, Xiang Zhang 0008, Long Lan, Xinwang Liu 0002, Kenli Li 0001, Xiaochun Cao
IEEE Trans. Circuits Syst. Video Technol.3
2026 Ultra-High-Definition Image Restoration: New Benchmarks and a Dual Interaction Prior-Driven Solution
abstract
Ultra-High-Definition (UHD) image restoration has acquired remarkable attention due to its practical demand. In this paper, we construct UHD snow and rain benchmarks, named UHD-Snow and UHD-Rain, to remedy the deficiency in this field. The UHD-Snow/UHD-Rain is established by simulating the physics process of rain/snow into consideration and each benchmark contains 3200 degraded/clear image pairs of 4K resolution. Furthermore, we propose an effective UHD image restoration solution by considering gradient and normal priors in model design, thanks to these priors’ spatial and detail contributions. Specifically, our method contains two branches: (a) feature fusion and reconstruction branch in high-resolution space and (b) prior feature interaction branch in low-resolution space. The former learns high-resolution features and fuses prior-guided low-resolution features to reconstruct clear images, while the latter utilizes normal and gradient priors to mine useful spatial features and detail features to guide high-resolution recovery better. To better utilize these priors, we introduce single prior feature interaction and dual prior feature interaction, where the former respectively fuses normal and gradient priors with high-resolution features to enhance prior ones, while the latter calculates the similarity between enhanced prior ones and further exploits dual guided filtering to boost the feature interaction of dual priors. We conduct experiments on both new and existing public datasets and demonstrate the state-of-the-art performance of our method on UHD image low-light enhancement, dehazing, deblurring, desnowing, and deraining. The source codes and benchmarks are available at https://github.com/wlydlut/UHDDIP.
Cong Wang 0018, Jinshan Pan, Xiaofeng Liu 0001, Weixiang Zhou, Xiaoran Sun, Wei Wang 0335, Zhixun Su
IEEE Trans. Circuits Syst. Video Technol.2
2026 Degradation-Aware Prompt Learning With Cross-Modal Compensation for Adverse Weather Removal
abstract
Adverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one restoration models attempt to address multiple degradation types within a unified framework, but often lack explicit spatial and semantic modeling of degradation characteristics, limiting their adaptability to diverse weather conditions. To address this limitation, we propose a Degradation-Aware Cross-Modal Prompt Compensation Network (DCMPC-Net) that leverages cross-modal degradation cues from a pre-trained vision-language model to condition restoration features within a unified backbone. Specifically, our DCMPC-Net mainly consists of the Cross-Modal Prompt Generator (CMPG), Prompt-Guided Attention Alignment Module (PGAAM), and Dual Feature Compensation Module (DFCM). The CMPG integrates textual embeddings with visual features to produce degradation-aware prompts that encode degradation-related semantic and contextual cues. These prompts are injected into the decoder via a PGAAM, which adaptively aligns semantic information with degraded regions to facilitate context-aware restoration. To further enhance structural fidelity, DFCM is introduced that disentangles degradation artifacts from scene structures, thereby improving the reconstruction of fine textures and detailed content. By integrating cross-modal semantic guidance with spatial alignment and structural enhancement, DCMPC-Net achieves robust and perceptually consistent restoration across diverse weather conditions. Extensive experiments show that DCMPC-Net outperforms state-of-the-art methods in both task-specific and unified settings, achieving superior accuracy and visual fidelity. The code is available at https://github.com/fanamber831/DCMPC-Net.
Wanshu Fan, Yunzhe Zhang, Jing Qin 0007, Kin-Man Lam 0001, Cong Wang 0018, Jinshan Pan
IEEE Trans. Image Process.7
2026 Multi-Modal Object Re-Identification With Prompt-S6 and Semantic-Aware Knowledge Guidance
abstract
Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However, existing multi-modal ReID methods do not effectively address background interference suppression or achieve tri-modal alignment, instead focusing on pairwise feature fusion. Moreover, many current aggregation approaches suffer from high computational complexity. To address these limitations, we propose PRISM, a novel multi-modal ReID framework built upon Prompt-S6 (PS6) and semantic-aware knowledge guidance. PS6 maintains the linear complexity and strong sequence modeling capability of Mamba while enabling efficient cross-modal interaction. Leveraging these advantages, we design two key components: Semantic-Driven Token Pruning (SDTP) and Progressive Fusion Network (PFN). Parsing semantic priors from the segmentation foundation models, the SDTP then leverages these priors and applies dynamic token pruning to suppress background noise and refine feature representations. The PFN progressively aggregates multi-modal features to achieve tri-modal alignment and fully exploit modality complementarity. With the proposed modules, PRISM generates more robust multi-modal representations under complex scenarios. Extensive experiments on four multi-modal object ReID benchmarks demonstrate the effectiveness and efficiency of our approach. The source code is available at https://github.com/zw-absin/PRISM.
Weixiang Zhou, Jiabei Zuo, Cong Wang 0018, Huchuan Lu, Zhixun Su
IEEE Trans. Image Process.4
2026 Federated Learning-Based Distributed Data Completion in Sparse Mobile CrowdSensing
abstract
Sparse Mobile CrowdSensing (SMCS) is an emerging distributed data collection framework. As one of the core methods,data completion uses the collected data to fill in the missing data. However, this approach inevitably poses significant privacy risks, since the traditional data completion methods require users' time and location information. In this paper, we propose a federated learning-based distributed data completion framework, which employs matrix factorization (MF) for local completion model training and federated learning to aggregate parameters of the MF model. This enables the construction of a global completion model without requiring private data, thereby mitigating privacy concerns. To address the challenges posed by the sparsity and asynchrony of distributed data, we incorporate time-aware deep neural network architectures with an asynchronous mechanism to enable federated training under irregularly gathered local data. Experimental results demonstrate that the proposed model achieves high data completion accuracy while ensuring robust privacy protection.
En Wang, Baoju Li, Zengyi Han, Cong Wang 0018, Ximing Li 0002, Jie Wu 0001
IEEE Trans. Mob. Comput.5
2025 Intra and Inter Parser-Prompted Transformers for Effective Image Restoration
abstract
We propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for restoring images from degraded observations and a Parser-Prompted Feature Generation Network (PPFGNet) for providing IRNet with reliable parser information to boost restoration. To enhance the integration of the parser within IRNet, we propose Intra Parser-Prompted Attention (IntraPPA) and Inter Parser-Prompted Attention (InterPPA) to implicitly and explicitly learn useful parser features to facilitate restoration. The IntraPPA re-considers cross attention between parser and restoration features, enabling implicit perception of the parser from a long-range and intra-layer perspective. Conversely, the InterPPA initially fuses restoration features with those of the parser, followed by formulating these fused features within an attention mechanism to explicitly perceive parser information. Further, we propose a parser-prompted feed-forward network to guide restoration within pixel-wise gating modulation. Experimental results show that PPTformer achieves state-of-the-art performance on image deraining, defocus deblurring, desnowing, and low-light enhancement.
Cong Wang 0018, Jinshan Pan, Wei Wang 0335
AAAI1
2025 DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
abstract
Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods frequently integrate semantic information from images or simply concatenate images, which often leads to low fidelity and flickering in the generated videos. To tackle these problems, we propose a high-fidelity image-to-video generation method by devising a frame retention branch based on a pre-trained video diffusion model, named DreamVideo. Our DreamVideo perceives the reference image via convolution layers and concatenates the features with the noisy latents as model input. By this means, the details of the reference image can be preserved to the greatest extent. In addition, by incorporating the designed double-condition classifier-free guidance, DreamVideo can generate high-quality videos of different actions by providing varying prompt texts. We conduct comprehensive experiments on the public datasets, and both quantitative and qualitative results indicate that our method outperforms the state-of-the-art method.
Cong Wang 0018, Jiaxi Gu, Panwen Hu, Yuanfan Guo, Hang Xu 0004, Xiaodan Liang
ICASSP1
2025 Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
abstract
Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the Motion-priors Conditional Diffusion Model (MCDM), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also introduce the TalkingFace-Wild dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation.
Fei Shen 0004, Cong Wang 0018, Junyao Gao 0002, Jisheng Dang, Jinhui Tang 0001, Tat-Seng Chua
ICML2
2025 Rethinking Joint Maximum Mean Discrepancy for Visual Domain Adaptation
abstract
In domain adaption (DA), joint maximum mean discrepancy (JMMD), as a famous distribution-distance metric, aims to measure joint probability distribution difference between the source domain and target domain, while it is still not fully explored and especially hard to be applied into a subspace-learning framework as its empirical estimation involves a tensor-product operator whose partial derivative is difficult to obtain. To solve this issue, we deduce a concise JMMD based on the Representer theorem that avoids the tensor-product operator and obtains two essential findings. First, we reveal the uniformity of JMMD by proving that previous marginal, class conditional, and weighted class conditional probability distribution distances are three special cases of JMMD with different label reproducing kernels. Second, inspired by graph embedding, we observe that the similarity weights, which strengthen the intra-class compactness in the graph of Hilbert Schmidt independence criterion (HSIC), take opposite signs in the graph of JMMD, revealing why JMMD degrades the feature discrimination. This motivates us to propose a novel loss JMMD-HSIC by jointly considering JMMD and HSIC to promote discrimination of JMMD. Extensive experiments on several cross-domain datasets could demonstrate the validity of our revealed theoretical results and the effectiveness of our proposed JMMD-HSIC.
Wei Wang 0335, Haifeng Xia, Chao Huang 0008, Zhengming Ding, Cong Wang 0018, Xiaochun Cao
NeurIPS5
2025 GTIGNet: Global Topology Interaction Graphormer Network for 3D hand pose estimation
Wanshu Fan, Cong Wang 0018, Shixi Wen, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
Neural Networks3
2025 FDC: Feature Dropout Consistency for unsupervised domain adaptation semantic segmentation
Chaoyu Rao, Wanshu Fan, Cong Wang 0018, Xin Yang 0011, Xiaopeng Wei
Neural Networks3
2025 UniAdapter: All-in-One Control for Flexible Video Generation
abstract
Condition-based video generation aims to create video content based on given information that describes specific subjects. However, most existing works can only utilize a single condition to guide the denoising process, thereby limiting their applicability to specific scenarios. Although some works attempt to accommodate multiple conditions within one framework, they often require multiple encoders, leading to inefficiencies in integrating multi-condition features. In this work, we present a framework that, with the support of the proposed Unified Adapter (UniAdapter), enables simultaneous multi-condition control of video generation within a single model. To effectively merge these conditions, we propose a novel Probabilistic Multi-condition Concatenator (PMC) module, which employs a unified encoder to accommodate multiple conditions and concatenate condition features at the pixel level to achieve fine-grained control. Following the PMC module, we employ 2D down-sampling blocks to refine features for injection into the Video Diffusion Model (VDM). Moreover, our UniAdapter is designed to be model-agnostic and compatible with any U-Net-based VDM, offering a versatile solution for improving video generation quality. Experimental results on public benchmarks UCF-101 and MSR-VTT show that our method achieves superior results in both quantitative and qualitative evaluations.
Cong Wang 0018, Panwen Hu, Yuanfan Guo, Jiaxi Gu, Jianhua Han, Hang Xu 0004, Xiaodan Liang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Deep Label Propagation With Nuclear Norm Maximization for Visual Domain Adaptation
abstract
Domain adaptation aims to leverage abundant label information from a source domain to an unlabeled target domain with two different distributions. Existing methods usually rely on a classifier to generate high-quality pseudo-labels for the target domain, facilitating the learning of discriminative features. Label propagation (LP), as an effective classifier, propagates labels from the source domain to the target domain by designing a smooth function over a similarity graph, which represents structural relationships among data points in feature space. However, LP has not been thoroughly explored in deep neural network-based domain adaptation approaches. Additionally, the probability labels generated by LP are low-confident and LP is sensitive to class imbalance problem. To address these problems, we propose a novel approach for domain adaptation named deep label propagation with nuclear norm maximization (DLP-NNM). Specifically, we employ the constraint of nuclear norm maximization to enhance both label confidence and class diversity in LP and propose an efficient algorithm to solve the corresponding optimization problem. Subsequently, we utilize the proposed LP to guide the classifier layer in a deep discriminative adaptation network using the cross-entropy loss. As such, the network could produce more reliable predictions for the target domain, thereby facilitating more effective discriminative feature learning. Extensive experimental results on three cross-domain benchmark datasets demonstrate that the proposed DLP-NNM surpasses existing state-of-the-art domain adaptation approaches.
Wei Wang 0335, Cong Wang 0018, Chao Huang 0008, Zhengming Ding, Feiping Nie 0001, Xiaochun Cao
IEEE Trans. Image Process.3
2025 Optimal Graph Learning-Based Label Propagation for Cross-Domain Image Classification
abstract
Label propagation (LP) is a popular semi-supervised learning technique that propagates labels from a training dataset to a test one using a similarity graph, assuming that nearby samples should have similar labels. However, the recent cross-domain problem assumes that training (source domain) and test data sets (target domain) follow different distributions, which may unexpectedly degrade the performance of LP due to small similarity weights connecting the two domains. To address this problem, we propose optimal graph learning-based label propagation (OGL2P), which optimizes one cross-domain graph and two intra-domain graphs to connect the two domains and preserve domain-specific structures, respectively. During label propagation, the cross-domain graph draws two labels close if they are nearby in feature space and from different domains, while the intra-domain graph pulls two labels close if they are nearby in feature space and from the same domain. This makes label propagation more insensitive to cross-domain problems. During graph embedding, we optimize the three graphs using features and labels in the embedded subspace to extract locally discriminative and domain-invariant features and make the graph construction process robust to noise in the original feature space. Notably, as a more relaxed constraint, locally discriminative and domain-invariant can somewhat alleviate the contradiction between discriminability and domain-invariance. Finally, we conduct extensive experiments on five cross-domain image classification datasets to verify that OGL2P outperforms some state-of-the-art cross-domain approaches.
Wei Wang 0335, Mengzhu Wang, Chao Huang 0008, Cong Wang 0018, Jie Mu, Feiping Nie 0001, Xiaochun Cao
IEEE Trans. Image Process.4
2024 Correlation Matching Transformation Transformers for UHD Image Restoration
abstract
This paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution space and (b) learning in low-resolution space. The former learns multi-level high-resolution features and fuses low-high features and reconstructs the residual images, while the latter explores more representative features learning from the high-resolution ones to facilitate better restoration. To better improve feature representation in low-resolution space, we propose to build feature transformation from the high-resolution space to the low-resolution one. To that end, we propose two new modules: Dual-path Correlation Matching Transformation module (DualCMT) and Adaptive Channel Modulator (ACM). The DualCMT selects top C/r (r is greater or equal to 1 which controls the squeezing level) correlation channels from the max-pooling/mean-pooling high-resolution features to replace low-resolution ones in Transformers, which can effectively squeeze useless content to improve the feature representation in low-resolution space to facilitate better recovery. The ACM is exploited to adaptively modulate multi-level high-resolution features, enabling to provide more useful features to low-resolution space for better learning. Experimental results show that our UHDformer reduces about ninety-seven percent model sizes compared with most state-of-the-art methods while significantly improving performance under different training sets on 3 UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring. The source codes will be made available at https://github.com/supersupercong/UHDformer.
Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Mengzhu Wang, Xiao-Ming Wu 0003, Jun Liu 0036
AAAI1
2024 SelfPromer: Self-Prompt Dehazing Transformers with Depth-Consistency
abstract
This work presents an effective depth-consistency Self-Prompt Transformer, terms as SelfPromer, for image dehazing. It is motivated by an observation that the estimated depths of an image with haze residuals and its clear counterpart vary. Enforcing the depth consistency of dehazed images with clear ones, therefore, is essential for dehazing. For this purpose, we develop a prompt based on the features of depth differences between the hazy input images and corresponding clear counterparts that can guide dehazing models for better restoration. Specifically, we first apply deep features extracted from the input images to the depth difference features for generating the prompt that contains the haze residual information in the input. Then we propose a prompt embedding module that is designed to perceive the haze residuals, by linearly adding the prompt to the deep features. Further, we develop an effective prompt attention module to pay more attention to haze residuals for better removal. By incorporating the prompt, prompt embedding, and prompt attention into an encoder-decoder network based on VQGAN, we can achieve better perception quality. As the depths of clear images are not available at inference, and the dehazed images with one-time feed-forward execution may still contain a portion of haze residuals, we propose a new continuous self-prompt inference that can iteratively correct the dehazing model towards better haze-free image generation. Extensive experiments show that our SelfPromer performs favorably against the state-of-the-art approaches on both synthetic and real-world datasets in terms of perception metrics including NIQE, PI, and PIQE. The source codes will be made available at https://github.com/supersupercong/SelfPromer.
Cong Wang 0018, Jinshan Pan, Wanyu Lin, Jiangxin Dong, Wei Wang 0335, Xiao-Ming Wu 0003
AAAI1
2024 Sharpness-Aware Model-Agnostic Long-Tailed Domain Generalization
abstract
Domain Generalization (DG) aims to improve the generalization ability of models trained on a specific group of source domains, enabling them to perform well on new, unseen target domains. Recent studies have shown that methods that converge to smooth optima can enhance the generalization performance of supervised learning tasks such as classification. In this study, we examine the impact of smoothness-enhancing formulations on domain adversarial training, which combines task loss and adversarial loss objectives. Our approach leverages the fact that converging to a smooth minimum with respect to task loss can stabilize the task loss and lead to better performance on unseen domains. Furthermore, we recognize that the distribution of objects in the real world often follows a long-tailed class distribution, resulting in a mismatch between machine learning models and our expectations of their performance on all classes of datasets with long-tailed class distributions. To address this issue, we consider the domain generalization problem from the perspective of the long-tail distribution and propose using the maximum square loss to balance different classes which can improve model generalizability. Our method's effectiveness is demonstrated through comparisons with state-of-the-art methods on various domain generalization datasets. Code: https://github.com/bamboosir920/SAMALTDG.
Houcheng Su, Weihao Luo, Daixian Liu, Mengzhu Wang, Junyang Chen 0001, Cong Wang 0018, Zhenghan Chen
AAAI7
2024 Batch Singular Value Polarization and Weighted Semantic Augmentation for Universal Domain Adaptation
abstract
As a more challenging domain adaptation setting, universal domain adaptation (UniDA) introduces category shift on top of domain shift, which needs to identify unknown category in the target domain and avoid misclassifying target samples into source private categories. To this end, we propose a novel UniDA approach named Batch Singular value Polarization and Weighted Semantic Augmentation (BSP-WSA). Specifically, we adopt an adversarial classifier to identify the target unknown category and align feature distributions between the two domains. Then, we propose to perform SVD on the classifier's outputs to maximize larger singular values while minimizing those smaller ones, which could prevent target samples from being wrongly assigned to source private classes. To better bridge the domain gap, we propose a weighted semantic augmentation approach for UniDA to generate data on common categories between the two domains. Extensive experiments on three benchmarks demonstrate that BSP-WSA could outperform existing state-of-the-art UniDA approaches.
Wangzi Qi, Wei Wang 0169, Chao Huang 0008, Jie Wen 0001, Cong Wang 0018
ICML5
2024 Explore Internal and External Similarity for Single Image Deraining with Graph Neural Networks
Cong Wang 0018, Wei Wang 0335, Chengjin Yu, Jie Mu
IJCAI1
2024 Optimal Graph Learning and Nuclear Norm Maximization for Deep Cross-Domain Robust Label Propagation
Wei Wang 0335, Chao Huang 0008, Yang Cao 0011, Cong Wang 0018, Xiaochun Cao
IJCAI6
2024 Multi-Attention Based Visual-Semantic Interaction for Few-Shot Learning
Peng Zhao 0010, Jie Mu, Huiting Liu 0001, Cong Wang 0018, Xiaochun Cao
IJCAI6
2024 Progressive Local and Non-Local Interactive Networks with Deeply Discriminative Training for Image Deraining
abstract
In this paper, we develop a progressive local and non-local interactive network with multi-scale cross-content deeply discriminative learning to solve image deraining. The proposed model contains two key techniques: 1) Progressive Local and Non-Local Interactive Network (PLNLIN) and 2) Multi-Scale Cross-Content Deeply Discriminative Learning (MCDDL). The PLNLIN is a U-shaped encoder-decoder network, where the proposed new Progressive Local and Non-Local Interactive Module (PLNLIM) is the basic unit in the encoder-decoder framework. The PLNLIM fully explores local and non-local learning in convolution and Transformer operation respectively and the local and non-local content are further interactively learned in a progressive manner. The proposed MCDDL not only discriminates the output of the generator but also receives the deep content from the generator to distinguish real and fake features at each side layer of the discriminator in a multi-scale manner. We show that the proposed MCDDL has fast and stable convergence properties that lack in existing discriminative learning manners. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods on five public synthetic datasets and one real-world data. The source codes will be made available at https://github.com/supersupercong/PLNLIN-MCDDL.
Cong Wang 0018, Jie Mu, Chengjin Yu, Wei Wang 0335
ACM Multimedia1
2024 PercepLIE: A New Path to Perceptual Low-Light Image Enhancement
abstract
While current CNN-based low-light image enhancement (LIE) approaches have achieved significant progress, they often fail to generate better perceptual quality which requires restoring better details and more natural colors. To address these problems, we set a new path, called PercepLIE, by presenting the VQGAN with Multi-luminance Detail Compensation (MDC) and Global Color Adjustment (GCA). Specifically, observed that latent light features of the low-light images are quite different from those captured in normal light, we utilize VQGAN to explore the latent light representation of normal-light images to help the estimation of the low-light and normal-light mapping. Furthermore, we employ Gamma correction with varying Gamma values on the gradient to create multi-luminance details, forming the basis for our MDC module to facilitate better detail estimation. To optimize the colors of low-light input images, we introduce a simple yet effective GCA module that is based on spatially-varying representation between the estimated normal-light images in this module and low-light inputs. By combining the VQGAN with MDC and GCA within a stage-wise training mechanism, our method generates images with finer details and natural colors and achieves favorable performance on both synthetic and real-world datasets in terms of perceptual quality metrics including NIQE, PI, and LPIPS. The source codes will be made available at https://github.com/supersupercong/PercepLIE.
Cong Wang 0018, Chengjin Yu, Jie Mu, Wei Wang 0335
ACM Multimedia1
2024 Coarse-to-fine mechanisms mitigate diffusion limitations on image restoration
Qinyu Yang, Cong Wang 0018, Wei Wang 0335, Zhixun Su
Comput. Vis. Image Underst.3
2024 Estimating High-Resolution Surface Normals via Low-Resolution Photometric Stereo Images
abstract
Acquiring high-resolution 3D surface structures is a crucial task in computer vision as it provides more detailed surface textures and clearer structures. Photometric stereo can measure per-pixel surface normals of a 3D object using various shading cues. However, obtaining high-resolution images in a linear response photometric stereo imaging system can be challenging. Additionally, photometric stereo, as a per-pixel reconstruction method, requires higher-resolution surface normal maps to accurately depict complex surface structures, particularly in regions that demand more attention and precise reconstruction. Therefore, measuring high-resolution surface normals via low-resolution photometric stereo images is of great importance. Motivated by these, we propose a Super-resolution Photometric Stereo Network, namely SR-PSN. In order to address the issues of measuring the high-resolution surface normals from low-resolution photometric images, we mainly (1) apply a dual-position threshold normalization pre-processing scheme to effectively handle the spatially-varying reflectance of non-Lambertian surfaces, (2) adopt a local affinity feature module to learn the rich structural representation by explicitly revealing the neighbor relationships, (3) employ a parallel multi-scale feature extractor, which preserves high-resolution representations and deep feature extraction, and (4) propose a shared-weight regressor to handle the multi-scale features, to prevent the model collapsing into learning non-important features related to a certain fixed scale. Extensive ablation experiments validate the effectiveness of our proposed modules. Furthermore, quantitative experiments conducted on public benchmarks demonstrate that SR-PSN outperforms state-of-the-art calibrated photometric stereo methods. Notably, SR-PSN achieves superior results while utilizing photometric stereo images with only half the resolution of other methods. It effectively restores the structure of complex surfaces, producing a high-resolution normal map.
Yakun Ju, Muwei Jian, Cong Wang 0018, Junyu Dong, Kin-Man Lam 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 PromptRestorer: A Prompting Image Restoration Method with Degradation Perception
abstract
We show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely hindered. To address this, we propose a Prompting image Restorer, termed as PromptRestorer. Specifically, PromptRestorer contains two branches: a restoration branch and a prompting branch. The former is used to restore images, while the latter perceives degradation priors to prompt the restoration branch with reliable perceived content to guide the restoration process for better recovery. To better perceive the degradation which is extracted by a pre-trained model from given degradation observations, we propose a prompting degradation perception modulator, which adequately considers the characters of the self-attention mechanism and pixel-wise modulation, to better perceive the degradation priors from global and local perspectives. To control the propagation of the perceived content for the restoration branch, we propose gated degradation perception propagation, enabling the restoration branch to adaptively learn more useful features for better recovery. Extensive experimental results show that our PromptRestorer achieves state-of-the-art results on 4 image restoration tasks, including image deraining, deblurring, dehazing, and desnowing.
Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Jiangxin Dong, Mengzhu Wang, Yakun Ju, Junyang Chen 0001
NeurIPS1
2023 Class-specific and self-learning local manifold structure for domain adaptation
Wei Wang 0335, Mengzhu Wang, Long Lan, Quannan Zu, Xiang Zhang 0008, Cong Wang 0018
Pattern Recognit.7
2022 Online-Updated High-Order Collaborative Networks for Single Image Deraining
abstract
Single image deraining is an important and challenging task for some downstream artificial intelligence applications such as video surveillance and self-driving systems. Most of the existing deep-learning-based methods constrain the network to generate derained images but few of them explore features from intermediate layers, different levels, and different modules which are beneficial for rain streaks removal. In this paper, we propose a high-order collaborative network with multi-scale compact constraints and a bidirectional scale-content similarity mining module to exploit features from deep networks externally and internally for rain streaks removal. Externally, we design a deraining framework with three sub-networks trained in a collaborative manner, where the bottom network transmits intermediate features to the middle network which also receives shallower rainy features from the top network and sends back features to the bottom network. Internally, we enforce multi-scale compact constraints on the intermediate layers of deep networks to learn useful features via a Laplacian pyramid. Further, we develop a bidirectional scale-content similarity mining module to explore features at different scales in a down-to-up and up-to-down manner. To improve the model performance on real-world images, we propose an online-update learning approach, which uses real-world rainy images to fine-tune the network and update the deraining results in a self-supervised manner. Extensive experiments demonstrate that our proposed method performs favorably against eleven state-of-the-art methods on five public synthetic datasets and one real-world dataset.
Cong Wang 0018, Jinshan Pan, Xiao-Ming Wu 0003
AAAI1
2022 Single image rain removal using recurrent scale-guide networks
Cong Wang 0018, Honghe Zhu, Wanshu Fan, Xiao-Ming Wu 0003, Junyang Chen 0001
Neurocomputing1
2022 Semi-Supervised Image Deraining Using Knowledge Distillation
abstract
Image deraining has achieved considerable progress based on supervised learning with synthetic training pairs, but is usually limited in handling real-world rainy images. Although semi-supervised methods are suggested to exploit real-world rainy images when training deep deraining models, their performances are still notably inferior. To address this crucial issue, this work proposes a semi-supervised image deraining network with knowledge distillation (SSID-KD) for better exploiting real-world rainy images. In particular, the consistency of feature distribution of rain streaks extracted from synthetic and real-world rainy images is enforced by adopting knowledge distillation. Moreover, as for the backbone in SSID-KD, we propose the multi-scale feature fusion module and the pyramid fusion module to better extract deep features of rainy images. SSID-KD can relieve the problem of over-deraining or under-deraining for real-world rainy images, while it can keep comparable performance with supervised deraining methods on several benchmark datasets. Extensive experiments on both synthetic and real-world rainy images have validated that our SSID-KD not only can achieve better deraining results than existing semi-supervised deraining methods but also are quantitatively comparable with state-of-the-art supervised deraining methods. Benefiting from the well exploration of real-world rainy images, our SSID-KD can obtain more visually plausible deraining results. The source code and trained models are publicly available athttps://github.com/cuiyixin555/SSID-KD.
Cong Wang 0018, Dongwei Ren, Yunjin Chen, Pengfei Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Self-Training Enhanced: Network Embedding and Overlapping Community Detection With Adversarial Learning
abstract
Network embedding (NE) aims to encode the relations of vertices into a low-dimensional space. After NE, we can obtain the learned vectors of vertices that preserve the proximity of network structures for subsequent applications, e.g., vertex classification and link prediction. In existing NE models, they usually exploit the skip-gram with a negative sampling method to optimize their objective functions. Generally, this method learns the vertex representation only from the local connectivity of vertices (i.e., neighbors). However, there is a larger scope of vertex connectivity in real-world scenarios: a vertex may have multifaceted aspects and should belong to overlapping communities. Taking a social network as the overlapping example, a user may subscribe to the channels of politics, economy, and sports simultaneously, but the politics share more common attributes with the economy and less with the sports. In this article, we propose an adversarial learning approach (ACNE) for modeling overlapping communities of vertices. Specifically, we map the association between communities and vertices into an embedding space. Moreover, we take further research on enhancing our ACNE with the following two operations. First, in the initialization stage, we adopt a walking strategy with perception to obtain paths containing more possible boundary vertices to improve overlapping community detection. Then, after representation learning with ACNE, we use soft community assignments from a simple classifier as supervision to update the weights of ACNE. This self-training mechanism referred to as ACNE-ST can help ACNE to achieve better performance. Experimental results demonstrate that the proposed methods, including ACNE and ACNE-ST, can outperform the state-of-the-art models on the subsequent tasks of vertex classification and overlapping community detection.
Junyang Chen 0001, Zhiguo Gong, Jiqian Mo, Wei Wang 0077, Wei Wang 0335, Cong Wang 0018, Weiwen Liu, Kaishun Wu
IEEE Trans. Neural Networks Learn. Syst.6
2022 Adversarial Caching Training: Unsupervised Inductive Network Representation Learning on Large-Scale Graphs
abstract
Network representation learning (NRL) has far-reaching effects on data mining research, showing its importance in many real-world applications. NRL, also known as network embedding, aims at preserving graph structures in a low-dimensional space. These learned representations can be used for subsequent machine learning tasks, such as vertex classification, link prediction, and data visualization. Recently, graph convolutional network (GCN)-based models, e.g., GraphSAGE, have drawn a lot of attention for their success in inductive NRL. When conducting unsupervised learning on large-scale graphs, some of these models employ negative sampling (NS) for optimization, which encourages a target vertex to be close to its neighbors while being far from its negative samples. However, NS draws negative vertices through a random pattern or based on the degrees of vertices. Thus, the generated samples could be either highly relevant or completely unrelated to the target vertex. Moreover, as the training goes, the gradient of NS objective calculated with the inner product of the unrelated negative samples and the target vertex may become zero, which will lead to learning inferior representations. To address these problems, we propose an adversarial training method tailored for unsupervised inductive NRL on large networks. For efficiently keeping track of high-quality negative samples, we design a caching scheme with sampling and updating strategies that has a wide exploration of vertex proximity while considering training costs. Besides, the proposed method is adaptive to various existing GCN-based models without significantly complicating their optimization process. Extensive experiments show that our proposed method can achieve better performance compared with the state-of-the-art models.
Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Cong Wang 0018, Zhenghua Xu 0001, Jianming Lv, Xueliang Li 0002, Kaishun Wu, Weiwen Liu
IEEE Trans. Neural Networks Learn. Syst.4
2021 Dense Feature Pyramid Grids Network for Single Image Deraining
abstract
Rainy images degrade the visional performance that may bring down the accuracy of various applications. In this paper, we propose a novel densely connected network with Dense Feature Pyramid Grids Modules, called DFPGN, to solve the rain removal task. Specifically, in the proposed DFPG, there are five operations from different layers with various pathways and scales as the input of the current layer so that each layer can fuse various features from shallower and deeper ones to improve the deraining ability of the network. Extensive experiments on real and synthetic rainy images are conducted to demonstrate the proposed method achieves superior rain removal performance over state-of-the-art approaches.
Cong Wang 0018, Zhixun Su, Junyang Chen 0001
ICASSP2
2021 Single image rain streak removal via layer similarity prior
Wanshu Fan, Yutong Wu 0002, Cong Wang 0018
Appl. Intell.3
2021 Pyramid fully residual network for single image de-raining
Guangle Yao, Cong Wang 0018, Yutong Wu 0002
Neurocomputing2
2021 TAM: Targeted Analysis Model With Reinforcement Learning on Short Texts
abstract
Mining topics on social media (e.g., Twitter and Facebook) is an important task for various applications, such as hot topic discovery, advertising, and promotion activities. Topic modeling techniques are helpful to find out topics that people are talking about. However, current full-analysis models cannot perform well on a focused analysis task-find out all topics related to one particular area in short documents. One reason is that the targeted topic is usually sparse in the corpus of short texts. Another one is, during clustering, even minor errors may compound and render the model useless. This article studies these problems and proposes a targeted analysis model (TAM) with reinforcement learning (RL) to extract any specific topic in a given corpus and perform fine-grained topic generation. In this work, we design a reward function of RL to prevent the false propagation problem induced by Gibbs sampling during the clustering. We amend the targeted topic modeling techniques to the case of RL and use policy search combined with the Gibbs EM algorithm for parameter estimation. Metrics of F1 score and the proposed normalized mutual information-F1 are exploited for the evaluation of clustering and topic generation, respectively. Our experiments have demonstrated that TAM can outperform state-of-the-art models-specifically achieving 25.7% improvement on the F1 score for binary clustering on average.
Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Weiwen Liu, Cong Wang 0018
IEEE Trans. Neural Networks Learn. Syst.6
2021 Single image deraining via deep shared pyramid network
Cong Wang 0018, Xiaoying Xing, Guangle Yao, Zhixun Su
Vis. Comput.1
2020 Physical Model Guided Deep Image Deraining
abstract
Single image deraining is an urgent task because the degraded rainy image makes many computer vision systems fail to work, such as video surveillance and autonomous driving. So, deraining becomes important and an effective deraining algorithm is needed. In this paper, we propose a novel network based on physical model guided learning for single image deraining, which consists of three sub-networks: rain streaks network, rain-free network, and guide-learning network. The concatenation of rain streaks and rain-free image that are estimated by rain streaks network, rain-free network, respectively, is input to the guide-learning network to guide further learning and the direct sum of the two estimated images is constrained with the input rainy image based on the physical model of rainy image. Moreover, we further develop the Multi-Scale Residual Block (MSRB) to better utilize multi-scale information and it is proved to boost the deraining performance. Quantitative and qualitative experimental results demonstrate that the proposed method outperforms the state-of-the-art deraining methods. The source code will be available at https://supercong94.wixsite.com/supercong94.
Honghe Zhu, Cong Wang 0018, Zhixun Su, Guohui Zhao
ICME2
2020 Joint Self-Attention and Scale-Aggregation for Self-Calibrated Deraining Network
abstract
In the field of multimedia, single image deraining is a basic pre-processing work, which can greatly improve the visual effect of subsequent high-level tasks in rainy conditions. In this paper, we propose an effective algorithm, called JDNet, to solve the single image deraining problem and conduct the segmentation and detection task for applications. Specifically, considering the important information on multi-scale features, we propose a Scale-Aggregation module to learn the features with different scales. Simultaneously, Self-Attention module is introduced to match or outperform their convolutional counterparts, which allows the feature aggregation to adapt to each channel. Furthermore, to improve the basic convolutional feature transformation process of Convolutional Neural Networks (CNNs), Self-Calibrated convolution is applied to build long-range spatial and inter-channel dependencies around each spatial location that explicitly expand fields-of-view of each convolutional layer through internal communications and hence enriches the output features. By designing the Scale-Aggregation and Self-Attention modules with Self-Calibrated convolution skillfully, the proposed model has better deraining results both on real-world and synthetic datasets. Extensive experiments are conducted to demonstrate the superiority of our method compared with state-of-the-art methods. The source code will be available at https://supercong94.wixsite.com/supercong94.
Cong Wang 0018, Yutong Wu 0002, Zhixun Su, Junyang Chen 0001
ACM Multimedia1
2020 DCSFN: Deep Cross-scale Fusion Network for Single Image Rain Removal
abstract
Rain removal is an important but challenging computer vision task as rain streaks can severely degrade the visibility of images that may make other visions or multimedia tasks fail to work. Previous works mainly focused on feature extraction and processing or neural network structure, while the current rain removal methods can already achieve remarkable results, training based on single network structure without considering the cross-scale relationship may cause information drop-out. In this paper, we explore the cross-scale manner between networks and inner-scale fusion operation to solve the image rain removal task. Specifically, to learn features with different scales, we propose a multi-sub-networks structure, where these sub-networks are fused via a cross-scale manner by Gate Recurrent Unit to inner-learn and make full use of information at different scales in these sub-networks. Further, we design an inner-scale connection block to utilize the multi-scale information and features fusion way between different scales to improve rain representation ability and we introduce the dense block with skip connection to inner-connect these blocks. Experimental results on both synthetic and real-world datasets have demonstrated the superiority of our proposed method, which outperforms over the state-of-the-art methods. The source code will be available at https://supercong94.wixsite.com/supercong94.
Cong Wang 0018, Xiaoying Xing, Yutong Wu 0002, Zhixun Su, Junyang Chen 0001
ACM Multimedia1
2020 Inductive Document Representation Learning for Short Text Clustering
Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Wei Wang 0335, Weiwen Liu, Cong Wang 0018
ECML/PKDD (3)7
2020 Single image deraining via nonlocal squeeze-and-excitation enhancing network
Cong Wang 0018, Wanshu Fan, Honghe Zhu, Zhixun Su
Appl. Intell.1
2020 Single image deraining via deep pyramid network with spatial contextual information aggregation
Cong Wang 0018, Yutong Wu 0002, Yu Cai 0004, Guangle Yao, Zhixun Su
Appl. Intell.1
2020 Weakly supervised single image dehazing
Cong Wang 0018, Wanshu Fan, Yutong Wu 0002, Zhixun Su
J. Vis. Commun. Image Represent.1
2020 Densely connected multi-scale de-raining net
Cong Wang 0018, Zhixun Su, Guangle Yao
Multim. Tools Appl.1
2019 Learning a multi-level guided residual network for single image deraining
Cong Wang 0018, Zhixun Su, Yutong Wu 0002, Guangle Yao
Signal Process. Image Commun.1