VLDB 2026 Research / reviewers in the wild / expert
Chao Wang 0102
dblp:188/7759-102
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-1297-768XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RegionSLM: Region-aware Question Answering on Document ScreenshotsabstractReal-world document question-answering that relies on screenshots, such as bills and forms, requires evidence that is often spatially localised and visually cluttered. However, most Screenshot Language Models (SLMs) encode the entire page holistically and rely on implicit attention to ''find'' relevant content, which limits both accuracy and efficiency. We present RegionSLM, a region-aware SLM designed to explicitly connect the question to its supporting regions. RegionSLM has two key components: (1) a patch-relevance router that learns a query–region relevance distribution, enabling the model to produce a box-free relevance prior at inference; and (2) Relevance-Guided Region Pooling (RGRP), a query-conditional attention–pooling module that aggregates dense features into a small set of region tokens, which preserves grounding signals while reducing computational overhead. To support training and evaluation, we further curate ReDoc, a region-supervised corpus with 105k documents and 350k question-answer pairs, obtained via a question-guided two-step filtering procedure. Extensive experiments on 12 datasets demonstrate that explicitly learning query–region relevance and pooling it into compact region tokens is an effective and practical recipe for document retrieval and understanding. Chao Wang 0102, Hehe Fan, Huichen Yang, Sarvnaz Karimi, Lina Yao 0001, Yi Yang 0001 |
SIGIR | 1 |
| 2025 | Adapting Text-to-Image Generation with Feature Difference Instruction for Generic Image RestorationabstractDiffusion-based Text-to-Image (T2I) models have demonstrated significant potential in image restoration. However, existing models continue to grapple with challenges such as complex training and prompt design. We introduce a new perspective for improving image restoration by injecting knowledge from pretrained vision-language models into current T2I models. We empirically show that the degradation and content representations in BLIP-2 can be linearly separated, providing promising degradation guidance for image restoration. Specifically, the Feature Difference Instruction (FDI) is first extracted by Q-Formers through a simple subtraction operation based on reference image pairs. Then, we propose a multi-scale FDI adapter to decouple the degradation style and corrupted artifacts, and inject the styleflow exclusively into specific blocks through adapter-tuning, thereby preventing noise interference and eschewing the need for cumbersome weight retraining. In this way, we can train various task-specific adapters according to different degradations, achieving rich detail enhancement in the restoration results. Furthermore, the proposed FDI adapters have attractive properties of practical value, such as composability and generalization ability for all-in-one and mixed-degradation restoration. Extensive experiments under various settings demonstrate that our method has promising repairing quality over 10 image restoration tasks and a wide range of other applications. Chao Wang 0102, Hehe Fan, Huichen Yang, Sarvnaz Karimi, Lina Yao 0001, Yi Yang 0001 |
CVPR | 1 |
| 2025 | ProtChatGPT: Towards Understanding Proteins with Hybrid Representation and Large Language ModelsabstractProtein research is crucial in various scientific disciplines, but understanding their intricate structure-function relationships remains challenging. Recent advancements in Large Language Models (LLMs) have significantly improved the comprehension of task-specific knowledge, suggesting the potential for specialized ChatGPT-like systems in protein research to aid fundamental investigations. In this work, we introduce ProtChatGPT, which aims to learn and understand protein structures using natural language. ProtChatGPT enables users to upload proteins, ask questions, and engage in interactive conversations to produce comprehensive answers. The system comprises multi-level protein encoding, protein-language alignment, and instruction tuning of LLMs. A protein first undergoes multiple protein encoders and PLP-former to produce multi-level hybrid protein embeddings, which are then aligned through a Protein Context Gating (PCG) module with contrastive learning, and projected by an adapter to conform with the LLM. The LLM finally combines user questions with projected protein embeddings to generate informative answers. Experiments show that ProtChatGPT can produce promising responses to proteins and the corresponding user questions. We hope that ProtChatGPT could form the basis for further exploration and application in protein research. Code and our pre-trained model will be publicly available. Chao Wang 0102, Hehe Fan, Ruijie Quan, Lina Yao 0001, Yi Yang 0001 |
SIGIR | 1 |
| 2025 | Training-free prior guided diffusion model for zero-reference low-light image enhancement
Kai Shang 0001, Ming-Wen Shao, Chao Wang 0102, Yuanjian Qiao 0001, Yecong Wan |
Neurocomputing | 3 |
| 2025 | SAMControl: Controlling Pose and Object for Image Editing with Soft Attention MaskabstractTo achieve content-consistent results in text-conditioned image editing, existing methods typically employ a reconstruction branch to capture the source image details via diffusion inversion and a generation branch to synthesize the target image based on the given textual prompt and the masked source image details. However, accurately segmenting source details is challenging with the current fixed-threshold mask strategy. Additionally, the inadequacies in the inversion process can lead to insufficient retention of source details. In this article, we propose a method called SAMControl (Soft Attention Mask) to adaptively control the pose and object details for image editing. SAMControl dynamically learns flexible attention masks for different images at various diffusion steps. Furthermore, in the reconstruction branch, we utilize a direct inversion technique to ensure the fidelity of source details within SAM. Extensive qualitative and quantitative results demonstrate the effectiveness of the proposed method. Yue Zhang 0004, Chao Wang 0102, Yunzhi Zhuge, Hehe Fan, Xiaojun Chang, Cheng Deng 0002, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Multi-Domain Multi-Scale Diffusion Model for Low-Light Image EnhancementabstractDiffusion models have achieved remarkable progress in low-light image enhancement. However, there remain two practical limitations: (1) existing methods mainly focus on the spatial domain for the diffusion process, while neglecting the essential features in the frequency domain; (2) conventional patch-based sampling strategy inevitably leads to severe checkerboard artifacts due to the uneven overlapping. To address these limitations in one go, we propose a Multi-Domain Multi-Scale (MDMS) diffusion model for low-light image enhancement. In particular, we introduce a spatial-frequency fusion module to seamlessly integrates spatial and frequency information. By leveraging the Multi-Domain Learning (MDL) paradigm, our proposed model is endowed with the capability to adaptively facilitate noise distribution learning, thereby enhancing the quality of the generated images. Meanwhile, we propose a Multi-Scale Sampling (MSS) strategy that follows a divide-ensemble manner by merging the restored patches under different resolutions. Such a multi-scale learning paradigm explicitly derives patch information from different granularities, thus leading to smoother boundaries. Furthermore, we empirically adopt the Bright Channel Prior (BCP) which indicates natural statistical regularity as an additional restoration guidance. Experimental results on LOL and LOLv2 datasets demonstrate that our method achieves state-of-the-art performance for the low-light image enhancement task. Codes are available at https://github.com/Oliiveralien/MDMS. Kai Shang 0001, Ming-Wen Shao, Chao Wang 0102, Yuanshuo Cheng, Shuigen Wang |
AAAI | 3 |
| 2024 | Depth-Aware Blind Image Decomposition for Real-World Adverse Weather Recovery
Chao Wang 0102, Zhedong Zheng, Ruijie Quan, Yi Yang 0001 |
ECCV (82) | 1 |
| 2024 | RDM-IR: Task-adaptive deep unfolding network for All-In-One image restoration
Yuanshuo Cheng, Ming-Wen Shao, Yecong Wan, Chao Wang 0102 |
Knowl. Based Syst. | 4 |
| 2024 | Multidimensional Dynamic Pruning: Exploring Spatial and Channel Fuzzy SparsityabstractDynamic pruning is an effective model compression method to reduce the computational cost of networks. However, existing dynamic pruning methods are limited to pruning along a single dimension (channel, spatial or depth), which cannot maximally excavate the redundancy of the network. Meanwhile, most of the current state-of-the-arts usually implement dynamic pruning via masked-out partial channels and pixels for training, while failing to accelerate the inference speed. To tackle these limitations, we propose a novel fuzzy-based Multi-Dimensional Dynamic Pruning (MDDP) paradigm to dynamically compress neural networks along both the channel and spatial dimensions. Specifically, we design a multi-dimensional fuzzy-mask block to simultaneously learn which spatial positions or channels are redundant and need to be pruned. Then, the Gumbel-Softmax trick combined with a sparsity loss is introduced to train these mask modules in an end-to-end manner. During the testing stage, we convert features and convolution kernels into two matrices respectively, and then implement sparse convolution through matrix multiplication to accelerate the network inference. Extensive experiments demonstrate that our method outperforms existing methods in terms of accuracy and computational cost. For instance, on the CIFAR-10 dataset, our method prunes 68% FLOPs of ResNet-56 with only a 0.07% Top-1 accuracy drop Ming-Wen Shao, Jiandong Kuang, Chao Wang 0102, Wangmeng Zuo, Guoyin Wang 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | Context-Aware Pretraining for Efficient Blind Image DecompositionabstractIn this paper, we study Blind Image Decomposition (BID), which is to uniformly remove multiple types of degradation at once without foreknowing the noise type. There remain two practical challenges: (1) Existing methods typically require massive data supervision, making them infeasible to real-world scenarios. (2) The conventional paradigm usually focuses on mining the abnormal pattern of a superimposed image to separate the noise, which de facto conflicts with the primary image restoration task. Therefore, such a pipeline compromises repairing efficiency and authenticity. In an attempt to solve the two challenges in one go, we propose an efficient and simplified paradigm, called Context-aware Pretraining (CP), with two pretext tasks: mixed image separation and masked image reconstruction. Such a paradigm reduces the annotation demands and explicitly facilitates context-aware feature learning. Assuming the restoration process follows a structure-to-texture manner, we also introduce a Context-aware Pretrained network (CPNet). In particular, CPNet contains two transformer-based parallel encoders, one information fusion module, and one multi-head prediction module. The information fusion module explicitly utilizes the mutual correlation in the spatial-channel dimension, while the multi-head prediction module facilitates texture-guided appearance flow. Moreover, a new sampling loss along with an attribute label constraint is also deployed to make use of the spatial context, leading to high-fidelity image restoration. Extensive experiments on both real and synthetic benchmarks show that our method achieves competitive performance for various BID tasks. Chao Wang 0102, Zhedong Zheng, Ruijie Quan, Yifan Sun 0003, Yi Yang 0001 |
CVPR | 1 |
| 2023 | FDDN: frequency-guided network for single image dehazing
Haozhen Shen, Chao Wang 0102, Liang-Jian Deng, Liangtian He, Ming-Wen Shao, Deyu Meng |
Neural Comput. Appl. | 2 |
| 2023 | Image Super-Resolution Using a Simple Transformer Without Pretraining
Huan Liu 0012, Ming-Wen Shao, Chao Wang 0102, Feilong Cao |
Neural Process. Lett. | 3 |
| 2022 | Context-Based Multiscale Unified Network for Missing Data Reconstruction in Remote Sensing ImagesabstractMissing data reconstruction is a classical yet challenging problem in remote sensing images. Most current methods based on traditional convolutional neural network require supplementary data and can only handle one specific task. To address these limitations, we propose a novel generative adversarial network-based missing data reconstruction method in this letter, which is capable of various reconstruction tasks given only single source data as input. Two auxiliary patch-based discriminators are deployed to impose additional constraints on the local and global regions, respectively. In order to better fit the nature of remote sensing images, we introduce special convolutions and attention mechanism in a two-stage generator, thereby benefiting the tradeoff between accuracy and efficiency. Combining with perceptual and multiscale adversarial losses, the proposed model can produce coherent structure with better details. Qualitative and quantitative experiments demonstrate the uncompromising performance of the proposed model against multisource methods in generating visually plausible reconstruction results. Moreover, further exploration shows a promising way for the proposed model to utilize spatio-spectral-temporal information. The codes and models are available athttps://github.com/Oliiveralien/Inpainting-on-RSI. Ming-Wen Shao, Chao Wang 0102, Tianjun Wu, Deyu Meng, Jiancheng Luo |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Dual-Pyramidal Image Inpainting With Dynamic NormalizationabstractDeep autoencoder-based approaches have achieved significant improvements on restoring damaged images, yet they still suffer from artifacts due to the inadequate representation and inaccurate regularization of existing features. In this paper, we propose a dual-pyramidal inpainting framework called DPNet to address these two limitations, which seamlessly integrates sufficient feature learning and dynamic regularization within an autoencoder network. Specifically, to exhaustively extract multi-scale features, we adopt layer-wise pyramidal convolution in encoder, which provides an arbitrary combination pool of various receptive fields. Subsequently, to tackle the patch deterioration problem in previous cross-scale non-local schemes, we further propose a Pyramidal Attention Mechanism (PAM) in decoder to acquire finer patches directly from learned layers. Mutually benefited with pyramidal features extraction in encoder, the dissemination space for non-local pixels in our PAM is notably enlarged to pyramidal level, thus significantly benefiting the feature representation. Moreover, to avoid the mask error accumulation in existing works, a dynamic normalization mechanism utilizing the spatial mask information updated in encoder is introduced, which further ensures the feature integrity and consistency. Such a dual-pyramidal structure along with dynamic normalization significantly improve the inpainting quality, outperforming existing competitors. Comprehensive experiments conducted on three benchmark datasets demonstrate that our DPNet performs favorably against the state-of-the-arts. Chao Wang 0102, Ming-Wen Shao, Deyu Meng, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Efficient Pyramidal GAN for Versatile Missing Data Reconstruction in Remote Sensing ImagesabstractMissing data reconstruction is a classical yet challenging problem in remote sensing image processing due to the complex atmospheric environment and variability of satellite sensors. Most of the contemporary reconstruction methods either handle only one specific task or require supplementary data, while the single-input for multi-task reconstruction has not been explored yet. In this paper we propose a novel Generative Adversarial Network-based unified framework for missing remote sensing image reconstruction, which is capable of various reconstruction tasks given only single source data as input. Specifically, we first propose a Mask Extraction Network (MEN) to obtain a united soft mask, which represents the intrinsic prior under various scenarios and indicates not only location but context information. The versatility of mask extraction enables the multi-task reconstruction of remote sensing images. Besides, we propose a Unified Inpainting Network (UIN) to repair diverse degraded images. Being specifically tailored for remote sensing images, Dilated pyramidal convolutions (DPC) and an Attention Fusion Mechanism (AFM) are introduced to further improve the feature extraction ability and thus exhaustly leveraging the single-input information. Extensive experiments demonstrate the uncompromising performance of the proposed method against state-of-the-art multi-input methods on diverse missing restoration. Moreover, further exploration shows the potential of the proposed method to utilize joint spatio-spectral-temporal information, which is evaluated to outperform existing competitors on remote sense images. Ming-Wen Shao, Chao Wang 0102, Wangmeng Zuo, Deyu Meng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | DMDIT: Diverse multi-domain image-to-image translation
Ming-Wen Shao, Youcai Zhang, Huan Liu 0012, Chao Wang 0102, Xun Shao |
Knowl. Based Syst. | 4 |