Chen Wu 0006

dblp:78/3213-6 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
17since 2021 · last 2026
0009-0002-1740-0804ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DCA-LUT: Deep Chromatic Alignment with 5D LUT for Purple Fringing Removal
abstract
Purple fringing, a persistent artifact caused by Longitudinal Chromatic Aberration (LCA) in camera lenses, has long degraded the clarity and realism of digital imaging. Traditional solutions rely on complex and expensive apochromatic (APO) lens hardware and the extraction of handcrafted features, ignoring the data-driven approach. To fill this gap, we introduce DCA-LUT, the first deep learning framework for purple fringing removal. Inspired by the physical root of the problem-the spatial misalignment of RGB color channels due to lens dispersion, we introduce a novel Chromatic-Aware Coordinate Transformation (CA-CT) module, learning an image-adaptive color space to decouple and isolate fringing into a dedicated dimension. This targeted separation allows the network to learn a precise "purple fringe channel," which then guides the accurate restoration of the luminance channel. The final color correction is performed by a learned 5D Look-Up Table (5D LUT), enabling efficient and powerful non-linear color mapping. To enable robust training and fair evaluation, we constructed a large-scale synthetic purple fringing dataset (PF-Synth). Extensive experiments in synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance in purple fringing removal.
Jialang Lu, Shuning Sun, Pu Wang 0008, Chen Wu 0006, Feng Gao 0005, Lina Gong, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
AAAI4
2026 CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare Removal
abstract
Purple flare, a diffuse chromatic aberration artifact commonly found around highlight areas, severely degrades the tone transition and color of the image. Existing traditional methods are based on hand-crafted features, which lack flexibility and rely entirely on fixed priors, while the scarcity of paired training data critically hampers deep learning. To address this issue, we propose a novel network built upon decoupled HSV Look-Up Tables (LUTs). The method aims to simplify color correction by adjusting the Hue (H), Saturation (S), and Value (V) components independently. This approach resolves the inherent color coupling problems in traditional methods. Our model adopts a two-stage architecture: First, a Chroma-Aware Spectral Tokenizer (CAST) converts the input image from RGB space to HSV space and independently encodes the Hue (H) and Value (V) channels into a set of semantic tokens describing the Purple flare status; second, the HSV-LUT module takes these tokens as input and dynamically generates independent correction curves (1D-LUTs) for the three channels H, S, and V. To effectively train and validate our model, we built the first large-scale purple flare dataset with diverse scenes. We also proposed new metrics and a loss function specifically designed for this task. Extensive experiments demonstrate that our model not only significantly outperforms existing methods in visual effects but also achieves state-of-the-art performance on all quantitative metrics.
Pu Wang 0008, Shuning Sun, Jialang Lu, Chen Wu 0006, Youshan Zhang, Chenggang Shan, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
AAAI4
2026 Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
Chen Wu 0006, Zhuoran Zheng, Jingyuan Xia, Weidong Jiang
ISCAS1
2026 Fusion requires interaction: a hybrid Mamba-transformer architecture for deep interactive fusion of multi-modal images
Wenxiao Xu, Chen Wu 0006, Qiyuan Yin, Zhuoran Zheng, Daqing Huang
Expert Syst. Appl.2
2026 Ultra-High-Definition Image Restoration via High-Frequency Enhanced Transformer
abstract
Transformer-based architectures exhibit substantial promise in the realm of ultra-high-definition (UHD) image restoration (IR). Nevertheless, they encounter significant challenges in maintaining high-frequency (HF) details, which are crucial for the reconstruction of texture. Conventional methods tackle computational complexity by significantly reducing the resolution (by a factor of 4 to 8). Moreover, the majority of high-frequency components are eliminated due to the inherent characteristics of self-attention mechanisms, as these mechanisms tend to naturally suppress high-frequency elements during non-local feature integration. This paper proposes a dual-branch transformer architecture that synergistically combines native-resolution HF preservation with efficient contextual modeling, named HiFormer. The high-resolution branch utilizes a directionally-sensitive large-kernel decomposition to effectively address anisotropic degradations with fewer parameters and applies depthwise separable convolutions for localized high-frequency (HF) information extraction. Concurrently, the low-resolution branch assimilates these localized HF elements using adaptive channel modulation to offset spectral losses induced by the inherent smoothing effect of self-attention. Comprehensive experiments across numerous UHD image restoration tasks reveal that our approach surpasses current leading methods in both quantitative metrics and qualitative analysis. The code is available at https://github.com/5chen/HiFormer.
Chen Wu 0006, Zhuoran Zheng, Weidong Jiang, Yuning Cui 0001, Jingyuan Xia
IEEE Trans. Circuits Syst. Video Technol.1
2026 Towards Ultra-High-Definition Image Deraining: A Benchmark and an Efficient Method
abstract
Despite significant advancements in image deraining, most existing methods are carried out on low-resolution images, leaving their effectiveness on high-resolution images uncertain. This limitation becomes even more pronounced with the rise of ultra-high-definition (UHD) imaging. In this paper, we tackle the challenge of UHD image deraining and introduce 4K-Rain13 k, the first large-scale UHD image deraining dataset, featuring 13,000 paired images at 4 K resolution. Leveraging this dataset, we conduct a benchmark study on existing methods for processing UHD images. To better address this task, we propose UDR-Mixer, an efficient and effective architecture tailored for UHD image deraining. Our model comprises two key components: a spatial feature rearrangement layer, which captures long-range dependencies in UHD images, and a frequency feature modulation layer, which enhances high-fidelity image reconstruction. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods while maintaining lower model complexity. The source code and proposed dataset are available athttps://github.com/cschenxiang/UDR-Mixer.
Hongming Chen 0004, Xiang Chen 0015, Chen Wu 0006, Zhuoran Zheng, Jinshan Pan, Xianping Fu
IEEE Trans. Multim.3
2025 DAP-LED: Learning Degradation-Aware Priors with Clip for Joint Low-Light Enhancement and Deblurring
abstract
Autonomous vehicles and robots often struggle with reliable visual perception at night due to the low illumination and motion blur caused by the long exposure time of RGB cameras. Existing methods address this challenge by sequentially connecting the off-the-shelf pretrained lowlight enhancement and deblurring models. Unfortunately, these methods often lead to noticeable artifacts (e.g., color distortions) in the over-exposed regions or make it hardly possible to learn the motion cues of the dark regions. In this paper, we interestingly find vision-language models, e.g., Contrastive LanguageImage Pretraining (CLIP), can comprehensively perceive diverse degradation levels at night. In light of this, we propose a novel transformer-based joint learning framework, named DAP-LED, which can jointly achieve low-light enhancement and deblurring, benefiting downstream tasks, such as depth estimation, segmentation, and detection in the dark. The key insight is to leverage CLIP to adaptively learn the degradation levels from images at night. This subtly enables learning rich semantic information and visual representation for optimization of the joint tasks. To achieve this, we first introduce a CLIPguided cross-fusion module to obtain multi-scale patch-wise degradation heatmaps from the image embeddings. Then, the heatmaps are fused via the designed CLIP-enhanced transformer blocks to retain useful degradation information for effective model optimization. Experimental results show that, compared to existing methods, our DAP-LED achieves state-of-the-art performance in the dark. Meanwhile, the enhanced results are demonstrated to be effective for three downstream tasks. For demo and more results, please check the project page: https://vlislab22.github.io/dap-led/.
Chen Wu 0006
ICRA2
2025 UniFlowRestore: A General Video Restoration Framework via Flow Matching and Prompt Guidance
Shuning Sun, Yu Zhang 0296, Chen Wu 0006, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
ACM Multimedia3
2025 Dropout the High-Rate Downsampling: A Novel Design Paradigm for UHD Image Restoration
abstract
With the popularization of high-end mobile devices, Ultra-high-definition (UHD) images have become ubiquitous in our lives. The restoration of UHD images is a highly challenging problem due to the exaggerated pixel count, which often leads to memory overflow during processing. Existing methods either downsample UHD images at a high rate before processing or split them into multiple patches for separate processing. However, high-rate downsampling leads to significant information loss, while patch-based approaches inevitably introduce boundary artifacts. In this paper, we propose a novel design paradigm to solve the UHD image restoration problem, called D2Net. D2Net enables direct full-resolution inference on UHD images without the need for high-rate downsampling or dividing the images into several patches. Specifically, we ingeniously utilize the characteristics of the frequency domain to establish long-range dependencies of features. Taking into account the richer local patterns in UHD images, we also design a multi-scale convolutional group to capture local features. Additionally, during the decoding stage, we dynamically incorporate features from the encoding stage to reduce the flow of irrelevant information. Extensive experiments on three UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring, show that our model achieves better quantitative and qualitative results than state-of-the-art methods.
Chen Wu 0006, Long Peng 0003, Dianjie Lu, Zhuoran Zheng
WACV1
2025 MixNet: Efficient global modeling for ultra-high-definition image restoration
Chen Wu 0006, Shuning Sun, Yu Zhang 0296, Zhuoran Zheng
Neurocomputing1
2025 Re-examine all-in-one image restoration: A catastrophic forgetting perspective
Chen Wu 0006, Pu Wang 0008, Zhuoran Zheng
Pattern Recognit. Lett.1
2025 Adaptive Feature Selection Modulation Network for Efficient Image Super-Resolution
abstract
In the realm of image super-resolution, learning-based methods have made significant progress. However, limited computational resources still restrict their application. This prompts us to develop an efficient method for achieving effective image super-resolution. In this letter, we propose a novel adaptive feature selection modulation network (AFSMNet) tailored for efficient image super-resolution. Specifically, we design feature modulation blocks, which include the adaptive feature selection modulation (AFSM) module and the self-gating feed-forward network (SFN). The AFSM module dynamically computes the importance of each feature channel. For channels with differing levels of importance, we employ distinct processing strategies, thereby concentrating the computational resources of the network on the more critical features as much as possible. This approach facilitates the maintenance of a low computational cost without compromising performance. The SFN restricts the flow of irrelevant feature information within the network through a simple gating mechanism. In this way, our method achieves efficient and effective image super-resolution. Extensive experiment results show that the proposed method achieves a better trade-off between reconstruction performance and computational efficiency compared to the current state-of-the-art lightweight super-resolution methods.
Chen Wu 0006, Xin Su 0009, Zhuoran Zheng
IEEE Signal Process. Lett.1
2025 NSDSAM: Noise-Suppression-Driven SAM for Infrared Small Target Detection
abstract
Although Segment Anything Model (SAM) have recently achieved remarkable progress, their generalization capability in infrared small target detection remains limited due to the inherently high noise levels in infrared imagery. To preserve the generalization and noise suppression ability of the model, we propose an method called NSDSAM, a noise-suppression-driven approach that enhances SAM at both internal and external levels. Internally, we develop a Hybrid Adapter for suppressing the noise of feature maps, consisting of an MLP adapter and a self-attention adapter. The self-attention adapter first performs entropy-aware reconstruction of features from noisy inputs and employs a gating mechanism for soft-attention fusion, mitigating SAM’s sensitivity to noise. Externally, we design a Spatial-Frequency hybrid Module (SFHM) that jointly processes spatial and frequency domains to overcome the self-attention model’s bias toward low-frequency components, further strengthening the suppression of background clutter and noise. Extensive experiments on multiple infrared datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance in infrared small target detection. The project code is available upon acceptance.
Wenxiao Xu, Qiyuan Yin, Chen Wu 0006, Dianjie Lu, Guijuan Zhang, Zhuoran Zheng
IEEE Trans. Geosci. Remote. Sens.3
2024 Rethinking Image Deraining via Text-guided Detail Reconstruction
abstract
Image deraining aims to recover clean images from degradation caused by rain streaks or raindrops of varying intensities. Recently many learning-based approaches have been proposed and achieved promising performance. However, these methods either focus on network architecture design or solely introduce image-level prior to the model. In this paper, we introduce text prior assisting the model in image deraining, as text descriptions have high flexibility and scalability. Text prior provides a wealth of semantic information to help the model achieve more detailed restoration, rather than blindly extrapolating details lost in the degraded image. To this end, we propose a novel image deraining framework based on the transformer, named TGDeraining. Specifically, to incorporate text prior into the framework, we design the Text Prior Embedded Transformer Block (TETB). TETB allows for dynamic guidance of the attention map guided by the text descriptions, thus emphasizing the restoration of critical missing details. The text prior is also fed to the feed-forward network to transform features in a controlled manner. Extensive experimental results demonstrate the effectiveness of our method in restoring a clear image using text as reference information.
Chen Wu 0006, Zhuoran Zheng, Pengwen Dai, Chenggang Shan, Xiuyi Jia
ICME1
2024 Trusted re-weighting for label distribution learning
abstract
Label distribution learning (LDL) is a novel machine learning paradigm that aims to shift 0/1 labels into descriptive degrees to characterize the polysemy of instances. Since the description degree takes a value between 0 \ensuremath{\sim} 1, it is difficult for the annotator to accurately annotate each label. Therefore, the predictive ability of numerous LDL algorithms may be degraded by the presence of noise in the label space. To address this problem, we propose a novel stability-trust LDL framework that aims to reconstruct the feature space of an arbitrary LDL dataset by using feature decoupling and prototype guidance. Specifically, first, we use prototype learning to select reliable cluster centers (representative vectors of label distributions) to filter out a set of clean samples (with labeled noise) on the original dataset. Then, we decouple the feature space (eliminating correlations among features) by modeling a weight assigner that is learned on this clean sample set, thus assigning weights to each sample of the original dataset. Finally, all existing LDL algorithms can be trained on this new re-weighted dataset for the goal of robust modeling. In addition, we create a new image dataset to support the training and testing of compared models. Experimental results demonstrate that the proposed framework boosts the performance of the LDL algorithm on datasets with label noise.
Zhuoran Zheng, Chen Wu 0006, Yeying Jin, Xiuyi Jia
UAI2
2024 Polyp-DAM: Polyp Segmentation via Depth Anything Model
abstract
Recently, large models (Segment Anything model) came on the scene to provide a new baseline for polyp segmentation tasks. This demonstrates that large models with a sufficient image level prior can achieve promising performance on a given task. In this paper, we unfold a new perspective on polyp segmentation modeling by leveraging the Depth Anything Model (DAM) to provide depth prior to polyp segmentation models. Specifically, the input polyp image is first passed through a frozen DAM to generate a depth map. The depth map and the input polyp images are then concatenated and fed into a convolutional neural network with multiscale to generate segmented images. Extensive experimental results demonstrate the effectiveness of our method, and in addition, we observe that our method still performs well on images of polyps with noise.
Zhuoran Zheng, Chen Wu 0006, Yeying Jin, Xiuyi Jia
IEEE Signal Process. Lett.2
2023 Dual-Domain Learning for JPEG Artifacts Removal
Guang Yang 0063, Chen Wu 0006, Feng Wang 0056
ICONIP (13)3
2020 Generating Well-Formed Answers by Machine Reading with Stochastic Selector Networks
Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003
AAAI2
2020 PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation
abstract
Self-supervised pre-training, such as BERT (Devlin et al., 2018), MASS (Song et al., 2019) and BART (Lewis et al., 2019), has emerged as a powerful technique for natural language understanding and generation.Existing pre-training techniques employ autoencoding and/or autoregressive objectives to train Transformer-based models by recovering original word tokens from corrupted text with some masked tokens.The training goals of existing techniques are often inconsistent with the goals of many language generation tasks, such as generative question answering and conversational response generation, for producing new text given context.This work presents PALM with a novel scheme that jointly pre-trains an autoencoding and autoregressive language model on a large unlabeled corpus, specifically designed for generating new text conditioned on context.The new scheme alleviates the mismatch introduced by the existing denoising scheme between pre-training and fine-tuning where generation is more than reconstructing original text.An extensive set of experiments show that PALM achieves new state-of-theart results on a variety of language generation benchmarks covering generative question answering (Rank 1 on the official MARCO leaderboard), abstractive summarization on CNN/DailyMail as well as Gigaword, question generation on SQuAD, and conversational response generation on Cornell Movie Dialogues.
Bin Bi, Chenliang Li 0003, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si
EMNLP (1)3
2020 StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
Wei Wang 0225, Bin Bi, Ming Yan 0008, Chen Wu 0006, Jiangnan Xia, Zuyi Bao, Liwei Peng, Luo Si
ICLR4
2019 A Deep Cascade Model for Multi-Document Reading Comprehension
abstract
A fundamental trade-off between effectiveness and efficiency needs to be balanced when designing an online question answering system. Effectiveness comes from sophisticated functions such as extractive machine reading comprehension (MRC), while efficiency is obtained from improvements in preliminary retrieval components such as candidate document selection and paragraph ranking. Given the complexity of the real-world multi-document MRC scenario, it is difficult to jointly optimize both in an end-to-end system. To address this problem, we develop a novel deep cascade learning model, which progressively evolves from the documentlevel and paragraph-level ranking of candidate texts to more precise answer extraction with machine reading comprehension. Specifically, irrelevant documents and paragraphs are first filtered out with simple functions for efficiency consideration. Then we jointly train three modules on the remaining texts for better tracking the answer: the document extraction, the paragraph extraction and the answer extraction. Experiment results show that the proposed method outperforms the previous state-of-the-art methods on two large-scale multidocument benchmark datasets, i.e., TriviaQA and DuReader. In addition, our online system can stably serve typical scenarios with millions of daily requests in less than 50ms.
Ming Yan 0008, Jiangnan Xia, Chen Wu 0006, Bin Bi, Zhongzhou Zhao, Ji Zhang 0011, Luo Si, Rui Wang 0005, Wei Wang 0225, Haiqing Chen
AAAI3
2019 Incorporating Relation Knowledge into Commonsense Reading Comprehension with Multi-task Learning
abstract
This paper focuses on how to take advantage of external relational knowledge to improve machine reading comprehension (MRC) with multi-task learning. Most of the traditional methods in MRC assume that the knowledge used to get the correct answer generally exists in the given documents. However, in real-world task, part of knowledge may not be mentioned and machines should be equipped with the ability to leverage external knowledge. In this paper, we integrate relational knowledge into MRC model for commonsense reasoning. Specifically, based on a pre-trained language model (LM), We design two auxiliary relation-aware tasks to predict if there exists any commonsense relation and what is the relation type be-tween two words, in order to better model the interactions between document and candidate answer option. We conduct experiments on two multi-choice benchmark datasets: the SemEval-2018 Task11 and the Cloze Story Test. The experimental results demonstrate the effectiveness of the proposed method, which achieves superior performance compared with the comparable baselines on both datasets.
Jiangnan Xia, Chen Wu 0006, Ming Yan 0008
CIKM2
2019 Incorporating External Knowledge into Machine Reading for Generative Question Answering
abstract
Bin Bi, Chen Wu, Ming Yan, Wei Wang, Jiangnan Xia, Chenliang Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003
EMNLP/IJCNLP (1)2
2018 Multi-Granularity Hierarchical Attention Fusion Networks for Reading Comprehension and Question Answering
abstract
This paper describes a novel hierarchical attention network for reading comprehension style question answering, which aims to answer questions for a given narrative paragraph.In the proposed method, attention and fusion are conducted horizontally and vertically across layers at different levels of granularity between question and paragraph.Specifically, it first encode the question and paragraph with fine-grained language embeddings, to better capture the respective representations at semantic level.Then it proposes a multi-granularity fusion approach to fully fuse information from both global and attended representations.Finally, it introduces a hierarchical attention network to focuses on the answer span progressively with multi-level softalignment.Extensive experiments on the large-scale SQuAD and TriviaQA datasets validate the effectiveness of the proposed method.At the time of writing the paper (Jan.12th 2018), our model achieves the first position on the SQuAD leaderboard for both single and ensemble models.We also achieves state-of-the-art results on TriviaQA, AddSent and AddOne-Sent datasets.
Wei Wang 0225, Chen Wu 0006, Ming Yan 0008
ACL (1)2
2017 Session-aware Information Embedding for E-commerce Product Recommendation
abstract
Most of the existing recommender systems assume that user's visiting history can be constantly recorded. However, in recent online services, the user identification may be usually unknown and only limited online user behaviors can be used. It is of great importance to model the temporal online user behaviors and conduct recommendation for the anonymous users. In this paper, we propose a list-wise deep neural network based architecture to model the limited user behaviors within each session. To train the model efficiently, we first design a session embedding method to pre-train a session representation, which incorporates different kinds of user search behaviors such as clicks and views. Based on the learnt session representation, we further propose a list-wise ranking model to generate the recommendation result for each anonymous user session. We conduct quantitative experiments on a recently published dataset from an e-commerce company. The evaluation results validate the effectiveness of the proposed method, which can outperform the state-of-the-art.
Chen Wu 0006, Ming Yan 0008
CIKM1