EDBT 2026 Demo / reviewers in the wild / expert
Guoli Wang 0004
dblp:30/7040-4
· DBLP profile ↗
23ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-8685-3968ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 11 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image RestorationabstractRecently, All-in-One image restoration (AiOIR) has advanced significantly, offering promising solutions for complex real-world degradations. However, most existing approaches heavily rely on degradation-specific representation learning, which can lead to oversmoothing and artifacts in the restored images. To address this limitation, we propose ClearAIR, a novel AiOIR framework inspired by human visual perception and designed with a hierarchical restoration strategy in a coarse-to-fine manner. First, leveraging the global priority characteristic of early human visual perception, we employ an image quality assessment model to evaluate the overall image structure and degradation level. Next, we introduce a Semantic Guidance Unit to provide coarse semantic region guidance and a Task Identifier to predict local degradation types, enabling a more informed characterization of local degradation patterns. Finally, aiming at the challenge of local detail restoration, we propose an Internal Clue Reuse Mechanism that deeply mines the internal information of the image in a self-supervised manner to enhance the model’s capacity for fine-detail recovery. Experimental results demonstrate that ClearAIR achieves superior restoration performance across diverse synthetic and real-world datasets. Xu Zhang 0044, Huan Zhang 0008, Guoli Wang 0004, Qian Zhang 0009, Lefei Zhang |
AAAI | 3 |
| 2026 | Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image RestorationabstractExisting All-in-One image restoration methods often fail to perceive degradation types and severity levels simultaneously, overlooking the importance of fine-grained quality perception. Moreover, these methods often utilize highly customized backbones, which hinder their adaptability and integration into more advanced restoration networks. To address these limitations, we propose Perceive-IR, a novel backbone-agnostic All-in-One image restoration framework designed for fine-grained quality control across various degradation types and severity levels. Its modular structure allows core components to function independently of specific backbones, enabling seamless integration into advanced restoration models without significant modifications. Specifically, Perceive-IR operates in two key stages: 1) multi-level quality-driven prompt learning stage, where a fine-grained quality perceiver is meticulously trained to discern three-tier quality levels by optimizing the alignment between prompts and images within the CLIP perception space. This stage ensures a nuanced understanding of image quality, laying the groundwork for subsequent restoration; 2) restoration stage, where the quality perceiver is seamlessly integrated with a difficulty-adaptive perceptual loss, forming a quality-aware learning strategy. This strategy not only dynamically differentiates sample learning difficulty but also achieves fine-grained quality control by driving the restored image toward the ground truth while pulling it away from both low- and medium-quality samples. Furthermore, Perceive-IR incorporates a Semantic Guidance Module (SGM) and Compact Feature Extraction (CFE). The SGM leverages semantic information from pre-trained vision models to provide high-level contextual guidance, while the CFE focuses on extracting degradation-specific features, ensuring accurate handling of diverse image degradations. Extensive experiments demonstrate that Perceive-IR not only surpasses state-of-the-art methods but also generalizes reliably to zero-shot real-world and unknown degraded scenes, while adapting seamlessly to different backbone networks. This versatility underscores the framework's robustness and backbone-agnostic design. Project page at https://house-yuyu.github.io/Perceive-IR/. Xu Zhang 0044, Jiaqi Ma 0002, Guoli Wang 0004, Qian Zhang 0009, Huan Zhang 0008, Lefei Zhang |
IEEE Trans. Image Process. | 3 |
| 2025 | UniUIR: Considering Underwater Image Restoration as an All-in-One LearnerabstractExisting underwater image restoration (UIR) methods generally only handle color distortion or jointly address color and haze issues, but they often overlook the more complex degradations that can occur in underwater scenes. To address this limitation, we propose a Universal Underwater Image Restoration method, termed as UniUIR, considering the complex scenario of real-world underwater mixed distortions as an all-in-one manner. To disentangle degradation-specific effects and capture their inter-correlations, we propose the Mamba Mixture-of-Experts module (MMoEM). Each expert specializes in distinct aspects of degradation, while gating mechanism dynamically routes features to appropriate experts. This design enables collaborative prior extraction and preserves global context, all within linear computational complexity. Building upon this foundation, to enhance degradation representation and address the task conflicts that arise when handling multiple types of degradation, we introduce the spatial-frequency prior generator. This module extracts degradation prior information in both spatial and frequency domains, and adaptively selects the most appropriate task-specific prompts based on image content, thereby improving the accuracy of image restoration. Finally, to more effectively address complex, region-dependent distortions in UIR task, we incorporate depth information derived from a large-scale pre-trained depth prediction model, thereby enabling the network to perceive and leverage depth variations across different image regions to handle localized degradation. Extensive experiments demonstrate that UniUIR can produce more attractive results across qualitative and quantitative comparisons, and shows strong generalization than state-of-the-art methods. Project page at https://house-yuyu.github.io/UniUIR. Xu Zhang 0044, Huan Zhang 0008, Guoli Wang 0004, Qian Zhang 0009, Lefei Zhang, Bo Du 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier DomainabstractRAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB images, a difference that goes beyond the color matrix and extends to spatial structure due to resolution variations. Recent methods directly rebuild color mapping and spatial structure via shared deep representation, limiting optimal performance. Inspired by Image Signal Processing (ISP) pipeline, which distinguishes image restoration and enhancement, we present a novel Neural ISP framework, named FourierISP. This approach breaks the image down into style and structure within the frequency domain, allowing for independent optimization. FourierISP is comprised of three subnetworks: Phase Enhance Subnet for structural refinement, Amplitude Refine Subnet for color learning, and Color Adaptation Subnet for blending them in a smooth manner. This approach sharpens both color and structure, and extensive evaluations across varied datasets confirm that our approach realizes state-of-the-art results. Code will be available at https://github.com/alexhe101/FourierISP. Xuanhua He, Tao Hu 0027, Guoli Wang 0004, Zejin Wang, Qian Zhang 0009, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
AAAI | 3 |
| 2024 | DiffTV: Identity-Preserved Thermal-to-Visible Face Translation via Feature Alignment and Dual-Stage ConditionsabstractThe thermal-to-visible (T2V) face translation task is essential for enabling face verification in low-light or dark conditions by converting thermal infrared faces into their visible counterparts. However, this task faces two primary challenges. First, the inherent differences between the modalities hinder the effective use of thermal information to guide RGB face reconstruction. Second, translated RGB faces often lack the identity details of the corresponding visible faces, such as skin color. To tackle these challenges, we introduce DiffTV, the first Latent Diffusion Model (LDM) specifically designed for T2V facial image translation with a focus on preserving identity. Our approach proposes a novel heterogeneous feature alignment strategy that bridges the modal gap and extracts both coarse-and fine-grained identity features consistent with visible images. Furthermore, a dual-stage condition injection strategy introduces control information to guide identity-preserved translation. Experimental results demonstrate the superior performance of DiffTV, particularly in scenarios where maintaining identity integrity is critical. Guiqin Zhao, Guoli Wang 0004, Zejin Wang, Antitza Dantcheva, Lan Du 0002, Cunjian Chen |
ACM Multimedia | 4 |
| 2024 | Eliminating and mining strategies for open-world object proposal
Cheng Wang 0048, Guoli Wang 0004, Qian Zhang 0009, Peng Guo 0001, Wenyu Liu 0001, Xinggang Wang |
Neurocomputing | 2 |
| 2024 | OpenInst: A simple query-based method for open-world instance segmentation
Cheng Wang 0048, Guoli Wang 0004, Qian Zhang 0009, Peng Guo 0001, Wenyu Liu 0001, Xinggang Wang |
Pattern Recognit. | 2 |
| 2023 | Restoration and enhancement on low exposure raw images by joint demosaicing and denoising
Jiaqi Ma 0002, Guoli Wang 0004, Lefei Zhang, Qian Zhang 0009 |
Neural Networks | 2 |
| 2022 | AdaptivePose: Human Parts as Adaptive PointsabstractMulti-person pose estimation methods generally follow top-down and bottom-up paradigms, both of which can be considered as two-stage approaches thus leading to the high computation cost and low efficiency. Towards a compact and efficient pipeline for multi-person pose estimation task, in this paper, we propose to represent the human parts as points and present a novel body representation, which leverages an adaptive point set including the human center and seven human-part related points to represent the human instance in a more fine-grained manner. The novel representation is more capable of capturing the various pose deformation and adaptively factorizes the long-range center-to-joint displacement thus delivers a single-stage differentiable network to more precisely regress multi-person pose, termed as AdaptivePose. For inference, our proposed network eliminates the grouping as well as refinements and only needs a single-step disentangling process to form multi-person pose. Without any bells and whistles, we achieve the best speed-accuracy trade-offs of 67.4% AP / 29.4 fps with DLA-34 and 71.3% AP / 9.1 fps with HRNet-W48 on COCO test-dev dataset. Yabo Xiao, Dongdong Yu, Guoli Wang 0004, Qian Zhang 0009, Mingshu He |
AAAI | 4 |
| 2022 | Learning Quality-Aware Representation for Multi-Person Pose RegressionabstractOff-the-shelf single-stage multi-person pose regression methods generally leverage the instance score (i.e., confidence of the instance localization) to indicate the pose quality for selecting the pose candidates. We consider that there are two gaps involved in existing paradigm: 1) The instance score is not well interrelated with the pose regression quality. 2) The instance feature representation, which is used for predicting the instance score, does not explicitly encode the structural pose information to predict the reasonable score that represents pose regression quality. To address the aforementioned issues, we propose to learn the pose regression quality-aware representation. Concretely, for the first gap, instead of using the previous instance confidence label (e.g., discrete {1,0} or Gaussian representation) to denote the position and confidence for person instance, we firstly introduce the Consistent Instance Representation (CIR) that unifies the pose regression quality score of instance and the confidence of background into a pixel-wise score map to calibrates the inconsistency between instance score and pose regression quality. To fill the second gap, we further present the Query Encoding Module (QEM) including the Keypoint Query Encoding (KQE) to encode the positional and semantic information for each keypoint and the Pose Query Encoding (PQE) which explicitly encodes the predicted structural pose information to better fit the Consistent Instance Representation (CIR). By using the proposed components, we significantly alleviate the above gaps. Our method outperforms previous single-stage regression-based even bottom-up methods and achieves the state-of-the-art result of 71.7 AP on MS COCO test-dev set. Yabo Xiao, Dongdong Yu, Guoli Wang 0004, Qian Zhang 0009 |
AAAI | 5 |
| 2022 | Cross-Domain Correlation Distillation for Unsupervised Domain Adaptation in Nighttime Semantic SegmentationabstractThe performance of nighttime semantic segmentation is restricted by the poor illumination and a lack of pixel-wise annotation, which severely limit its application in autonomous driving. Existing works, e.g., using the twilight as the intermediate target domain to perform the adaptation from daytime to nighttime, may fail to cope with the inherent difference between datasets caused by the camera equipment and the urban style. Faced with these two types of domain shifts, i.e., the illumination and the inherent difference of the datasets, we propose a novel domain adaptation framework via cross-domain correlation distillation, called CCDistill. The invariance of illumination or inherent difference between two images is fully explored so as to make up for the lack of labels for nighttime images. Specifically, we extract the content and style knowledge contained in features, calculate the degree of inherent or illumination difference between two images. The domain adaptation is achieved using the invariance of the same kind of difference. Extensive experiments on Dark Zurich and ACDC demon-strate that CCDistill achieves the state-of-the-art performance for nighttime semantic segmentation. Notably, our method is a one-stage domain adaptation network which can avoid affecting the inference time. Our implementation is available at https://github.com/ghuan99/CCDistill. Jichang Guo, Guoli Wang 0004, Qian Zhang 0009 |
CVPR | 3 |
| 2022 | ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative TransformerabstractIn order to get raw images of high quality for downstream Image Signal Process (ISP), in this paper we present an Efficient Locally Multiplicative Transformer called ELMformer for raw image restoration. ELMformer contains two core designs especially for raw images whose primitive attribute is single-channel. The first design is a Bi-directional Fusion Projection (BFP) module, where we consider both the color characteristics of raw images and spatial structure of single-channel. The second one is that we propose a Locally Multiplicative Self-Attention (L-MSA) scheme to effectively deliver information from the local space to relevant parts. ELMformer can efficiently reduce the computational consumption and perform well on raw image restoration tasks. Enhanced by these two core designs, ELMformer achieves the highest performance and keeps the lowest FLOPs on raw denoising and raw deblurring benchmarks compared with state-of-the-arts. Extensive experiments demonstrate the superiority and generalization ability of ELMformer. On SIDD benchmark, our method has even better denoising performance than ISP-based methods which need huge amount of additional sRGB training images. Jiaqi Ma 0002, Shengyuan Yan, Lefei Zhang, Guoli Wang 0004, Qian Zhang 0009 |
ACM Multimedia | 4 |
| 2022 | Deep momentum uncertainty hashing
Chaoyou Fu, Guoli Wang 0004, Xiang Wu 0001, Qian Zhang 0009, Ran He 0001 |
Pattern Recognit. | 2 |
| 2021 | Pseudo Facial Generation With Extreme Poses for Face RecognitionabstractFace recognition has achieved a great success in recent years, it is still challenging to recognize those facial images with extreme poses. Traditional methods consider it as a domain gap problem. Many of them settle it by generating fake frontal faces from extreme ones, whereas they are tough to maintain the identity information with high computational consumption and uncontrolled disturbances. Our experimental analysis shows a dramatic precision drop with extreme poses. Meanwhile, those extreme poses just exist minor visual differences after small rotations. Derived from this insight, we attempt to relieve such a huge precision drop by making minor changes to the input images without modifying existing discriminators. A novel lightweight pseudo facial generation is proposed to relieve the problem of extreme poses without generating any frontal facial image. It can depict the facial contour information and make appropriate modifications to preserve the critical identity information. Specifically, the proposed method reconstructs pseudo profile faces by minimizing the pixel-wise differences with original profile faces and maintaining the identity consistent information from their corresponding frontal faces simultaneously. The proposed framework can improve existing discriminators and obtain a great promotion on several benchmark datasets. Guoli Wang 0004, Jiaqi Ma 0002, Qian Zhang 0009, Jiwen Lu, Jie Zhou 0001 |
CVPR | 1 |
| 2021 | Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
Helong Zhou, Liangchen Song, Guoli Wang 0004, Junsong Yuan 0001, Qian Zhang 0009 |
ICLR | 5 |
| 2021 | Efficient recurrent attention network for remote sensing scene classificationabstractAbstract Scene classification for remote sensing is a popular topic, and many recent convolutional neural networks (CNNs)‐based methods have shown the great model capacity and learning ability of highly discriminative features. Given a large number of training data, CNN can extract extensive features and learn to predict a remote sensing image. However, for supervised learning tasks, deep models often rely on a large number of labelled remote sensing images, which are difficult to pre‐process. Thus, training a lightweight deep learning model is essential. Easy‐classified and hard samples may also cause an imbalance of training set and lead the model to overwhelm the loss function. Accordingly, a novel Efficient Recurrent Attention Network (ERANet) for remote sensing scene classification is proposed. Different from traditional deep learning methods, Efficientnet‐B0 is introduced as a lightweight backbone for the ARCNet framework, replacing the original one. By applying the modified efficient backbone, the low Floating Point Operations (FLOPs) and parameter numbers of the proposed ERANet are maintained. The significance of focal loss is determined and applied to address the sample imbalance problem and yield a desirable performance. Extensive experiments on several challenging remote sensing scene classification data sets prove the efficiency of the proposed ERANet. Le Liang, Guoli Wang 0004 |
IET Image Process. | 2 |
| 2021 | High-Fidelity Face Manipulation With Extreme Poses and ExpressionsabstractFace manipulation has shown remarkable advances with the flourish of Generative Adversarial Networks. However, due to the difficulties of controlling structures and textures, it is challenging to model poses and expressions simultaneously, especially for the extreme manipulation at high-resolution. In this article, we propose a novel framework that simplifies face manipulation into two correlated stages: a boundary prediction stage and a disentangled face synthesis stage. The first stage models poses and expressions jointly via boundary images. Specifically, a conditional encoder-decoder network is employed to predict the boundary image of the target face in a semi-supervised way. Pose and expression estimators are introduced to improve the prediction performance. In the second stage, the predicted boundary image and the input face image are encoded into the structure and the texture latent space by two encoder networks, respectively. A proxy network and a feature threshold loss are further imposed to disentangle the latent space. Furthermore, due to the lack of high-resolution face manipulation databases to verify the effectiveness of our method, we collect a new high-quality Multi-View Face (MVF-HQ) database. It contains 120,283 images at 6000 × 4000 resolution from 479 identities with diverse poses, expressions, and illuminations. MVF-HQ is much larger in scale and much higher in resolution than publicly available high-resolution face manipulation databases. We will release MVF-HQ soon to push forward the advance of face manipulation. Qualitative and quantitative experiments on four databases show that our method dramatically improves the synthesis quality. Chaoyou Fu, Yibo Hu 0001, Xiang Wu 0001, Guoli Wang 0004, Qian Zhang 0009, Ran He 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Multi-scale multi-patch person re-identification with exclusivity regularized softmax
Cheng Wang 0048, Liangchen Song, Guoli Wang 0004, Qian Zhang 0009, Xinggang Wang |
Neurocomputing | 3 |
| 2019 | Self-Ensembling Attention Networks: Addressing Domain Shift for Semantic SegmentationabstractRecent years have witnessed the great success of deep learning models in semantic segmentation. Nevertheless, these models may not generalize well to unseen image domains due to the phenomenon of domain shift. Since pixel-level annotations are laborious to collect, developing algorithms which can adapt labeled data from source domain to target domain is of great significance. To this end, we propose self-ensembling attention networks to reduce the domain gap between different datasets. To the best of our knowledge, the proposed method is the first attempt to introduce selfensembling model to domain adaptation for semantic segmentation, which provides a different view on how to learn domain-invariant features. Besides, since different regions in the image usually correspond to different levels of domain gap, we introduce the attention mechanism into the proposed framework to generate attention-aware features, which are further utilized to guide the calculation of consistency loss in the target domain. Experiments on two benchmark datasets demonstrate that the proposed framework can yield competitive performance compared with the state of the art methods. Yonghao Xu, Bo Du 0001, Lefei Zhang, Qian Zhang 0009, Guoli Wang 0004, Liangpei Zhang 0001 |
AAAI | 5 |
| 2019 | Neurons Merging Layer: Towards Progressive Redundancy Reduction for Deep Supervised HashingabstractDeep supervised hashing has become an active topic in information retrieval. It generates hashing bits by the output neurons of a deep hashing network. During binary discretization, there often exists much redundancy between hashing bits that degenerates retrieval performance in terms of both storage and accuracy. This paper proposes a simple yet effective Neurons Merging Layer (NMLayer) for deep supervised hashing. A graph is constructed to represent the redundancy relationship between hashing bits that is used to guide the learning of a hashing network. Specifically, it is dynamically learned by a novel mechanism defined in our active and frozen phases. According to the learned relationship, the NMLayer merges the redundant neurons together to balance the importance of each output neuron. Moreover, multiple NMLayers are progressively trained for a deep hashing network to learn a more compact hashing code from a long redundant code. Extensive experiments on four datasets demonstrate that our proposed method outperforms state-of-the-art hashing methods. Chaoyou Fu, Liangchen Song, Xiang Wu 0001, Guoli Wang 0004, Ran He 0001 |
IJCAI | 4 |
| 2017 | Ordinal pyramid coding for rotation invariant feature extraction
Guoli Wang 0004, Bin Fan 0001, Chunhong Pan |
Neurocomputing | 1 |
| 2017 | Feature Extraction by Rotation-Invariant Matrix Representation for Object Detection in Aerial ImageabstractThis letter proposes a novel rotation-invariant feature for object detection in optical remote sensing images. Different from previous rotation-invariant features, the proposed rotation-invariant matrix (RIM) can incorporate partial angular spatial information in addition to radial spatial information. Moreover, it can be further calculated between different rings for a redundant representation of the spatial layout. Based on the RIM, we further propose an RIM_FV_RPP feature for object detection. For an image region, we first densely extract RIM features from overlapping blocks; then, these RIM features are encoded into Fisher vectors; finally, a pyramid pooling strategy that hierarchically accumulates Fisher vectors in ring subregions is used to encode richer spatial information while maintaining rotation invariance. Both of the RIM and RIM_FV_RPP are rotation invariant. Experiments on airplane and car detection in optical remote sensing images demonstrate the superiority of our feature to the state of the art. Guoli Wang 0004, Xinchao Wang, Bin Fan 0001, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Ordinal pyramid pooling for rotation invariant object recognitionabstractLocal feature descriptor plays a fundamental role in many visual tasks, and its rotation invariance is a key issue for many recognition and detection problems. This paper proposes a novel rotation invariant descriptor by ordinal pyramid pooling of local Fourier transform features based on their radial gradient orientations. Since both the low-level feature and pooling strategy are rotation invariant, the obtained descriptor is rotation invariant by nature. Pooling based on orders of gradient orientations is not only invariant to in-plane rotation, but also encodes gradient orientation information into descriptor as well as spatial information to some extent. Moreover, these information is enhanced by the proposed pyramid pooling structure. Therefore, our method is naturally rotation invariant and has strong discriminative ability. Experimental results on the aerial car dataset demonstrate the effectiveness of our descriptor. Guoli Wang 0004, Bin Fan 0001, Chunhong Pan |
ICASSP | 1 |