EDBT 2026 Demo / reviewers in the wild / expert
Qi Zhu 0010
dblp:66/5923-10
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-1545-1854ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PulseMind: A Multi-Modal Medical Model for Real-World Clinical DiagnosisabstractRecent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To bridge this gap, we introduce PulseMind, a new family of multi-modal diagnostic models that integrates a systematically curated dataset, a comprehensive evaluation benchmark, and a tailored training framework. Specifically, we first construct a diagnostic dataset, MediScope, which comprises 98,000 real-world multi-turn consultations and 601,500 medical images, spanning over 10 major clinical departments and more than 200 sub-specialties. Then, to better reflect the requirements of real-world clinical diagnosis, we develop the PulseMind Benchmark, a multi-turn diagnostic consultation benchmark with a four-dimensional evaluation protocol comprising proactiveness, accuracy, usefulness, and language quality. Finally, we design a training framework tailored for multi-modal clinical diagnostics, centered around a core component named Comparison-based Reinforcement Policy Optimization (CRPO). Compared to absolute score rewards, CRPO uses relative preference signals from multi-dimensional comparisons to provide stable and human-aligned training guidance. Extensive experiments demonstrate that PulseMind achieves competitive performance on both the diagnostic consultation benchmark and public medical benchmarks. Jiangwei Lao, Qi Zhu 0010, Congyun Jin, Shinan Liu, Zhihong Lu 0002, Lihe Zhang, Jian Wang 0108 |
AAAI | 4 |
| 2025 | SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingabstractOpen-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two challenges. 1) Existing RS semantic categories are limited, particularly for pixel-level interpretation datasets. 2) Distinguishing among diverse RS spatial regions solely by language space is challenging due to the dense and intricate spatial distribution in open-world RS imagery. To address the first issue, we develop a fine-grained RS interpretation dataset, Sky-SA, which contains 183,375 high-quality local image-text pairs with full-pixel manual annotations, covering 1,763 category labels, exhibiting richer semantics and higher density than previous datasets. Afterwards, to solve the second issue, we introduce the vision-centric principle for vision-language modeling. Specifically, in the pre-training stage, the visual self-supervised paradigm is incorporated into image-text alignment, reducing the degradation of general visual representation capabilities of existing paradigms. Then, we construct a visual-relevance knowledge graph across open-category texts and further develop a novel vision-centric image-text contrastive loss for fine-tuning with text prompts. This new model, denoted as SkySense-O, demonstrates impressive zero-shot capabilities on a thorough evaluation encompassing 14 datasets over 4 tasks, from recognizing to reasoning and classification to localization. Specifically, it outperforms the latest models such as SegEarthOV, GeoRSCLIP, and VHM by a large margin, i.e., 11.95%, 8.04% and 3.55% on average respectively. The code is publicly available to facilitate further research at https://github.com/zqcrafts/SkySense-O. Qi Zhu 0010, Jiangwei Lao, Deyi Ji, Lixiang Ru, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Dong Liu 0002, Feng Zhao 0004 |
CVPR | 1 |
| 2025 | When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
Xue Yang 0005, Qi Zhu 0010, Jingdong Chen, Yansheng Li 0001 |
ICCV | 5 |
| 2025 | FourierMamba: Fourier Learning Integration with State Space Models for Image DerainingabstractImage deraining aims to remove rain streaks from rainy images and restore clear backgrounds. Currently, some research that employs the Fourier transform has proved to be effective for image deraining, due to it acting as an effective frequency prior for capturing rain streaks. However, despite there exists dependency of low frequency and high frequency in images, these Fourier-based methods rarely exploit the correlation of different frequencies for conjuncting their learning procedures, limiting the full utilization of frequency information for image deraining. Alternatively, the recently emerged Mamba technique depicts its effectiveness and efficiency for modeling correlation in various domains (e.g., spatial, temporal), and we argue that introducing Mamba into its unexplored Fourier spaces to correlate different frequencies would help improve image deraining. This motivates us to propose a new framework termed FourierMamba, which performs image deraining with Mamba in the Fourier space. Owing to the unique arrangement of frequency orders in Fourier space, the core of FourierMamba lies in the scanning encoding of different frequencies, where the low-high frequency order formats exhibit differently in the spatial dimension (unarranged in axis) and channel dimension (arranged in axis). Therefore, we design FourierMamba that correlates Fourier space information in the spatial and channel dimensions with distinct designs. Specifically, in the spatial dimension Fourier space, we introduce the zigzag coding to scan the frequencies to rearrange the orders from low to high frequencies, thereby orderly correlating the connections between frequencies; in the channel dimension Fourier space with arranged orders of frequencies in axis, we can directly use Mamba to perform frequency correlation and improve the channel information representation. Extensive experiments reveal that our method outperforms state-of-the-art methods both qualitatively and quantitatively. Dong Li 0055, Yidi Liu, Xueyang Fu, Jie Huang 0017, Senyan Xu, Qi Zhu 0010, Zhengjun Zha |
ICML | 6 |
| 2025 | ADMIRE: ADaptive method to enhance Multiple Image REsolutions in text-rich multi-image understanding
Qipeng Zhu, Zhihong Lu 0002, Jiangwei Lao, Congyun Jin, Yingzhe Peng, Qi Zhu 0010, Lianzhen Zhong, Jiajia Liu 0002, Jian Wang 0108 |
KDD (2) | 8 |
| 2024 | Empowering Resampling Operation for Ultra-High-Definition Image Enhancement with Model-Aware GuidanceabstractImage enhancement algorithms have made remarkable advancements in recent years, but directly applying them to Ultra-high-definition (UHD) images presents intractable computational overheads. Therefore, previous straightforward solutions employ resampling techniques to reduce the resolution by adopting a “Downsampling-Enhancement-Upsampling” processing paradigm. However, this paradigm disentangles the resampling operators and inner enhancement algorithms, which results in the loss of information that is favored by the model, further leading to sub-optimal outcomes. In this paper, we propose a novel method of Learning Model-Aware Resampling (LMAR), which learns to customize resampling by extracting model-aware information from the UHD input image, under the guidance of model knowledge. Specifically, our method consists of two core designs, namely compensatory kernel estimation and steganographic resampling. At the first stage, we dynamically predict compensatory kernels tailored to the specific input and resampling scales. At the second stage, the image-wise compensatory information is derived with the compensatory kernels and embedded into the rescaled input images. This promotes the representation of the newly derived downscaled inputs to be more consistent with the full-resolution UHD inputs, as perceived by the model. Our LMAR enables model-aware and model-favored resampling while maintaining compatibility with existing resampling operators. Extensive experiments on multiple UHD image enhancement datasets and different backbones have shown consistent performance gains after correlating resizer and enhancer; e.g., up to 1.2dB PSNR gain for ×1.8 resampling scale on UHD-LOL4K. The code is available at https://github.com/YPatrickW/LMAR. Jie Huang 0017, Bing Li 0024, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004 |
CVPR | 5 |
| 2024 | Learning Spatio-Temporal Sharpness Map for Video DeblurringabstractVideo deblurring is a challenging task because only input blurry sequences are available. To further constrain the optimization process, existing methods explore various additional information,e.g., events, depth and sharpness prior. However, they consume large computing costs or generate unpleasant visual results due to the insufficient exploitation of spatio-temporal information. In this work, we develop a novel spatio-temporal sharpness map learned by a prior-based generation network implicitly. The proposed generation network blends both spatial and temporal sharpness priors in a blurry sequence, while few extra parameters are added. We show that the proposed map has better spatial continuity and guidance for video deblurring than the previous method. Furthermore, different from the simply concatenation in the previous work, we allow the sharpness map to accommodate to more effective video deblurring via a dual-stream network. Specifically, the network is decomposed by two branches, namely inter-frame and intra-frame reconstructions. The inter-frame reconstruction obtains the sharp patches of cecutive frames from the sharpness map to restore textures well. Meanwhile, the other intra-frame branch is responsible for recovering structures of the latent frame, where a novel histogram statistical method is developed to quantify and count textures in the feature under the modulation of the sharpness map. Quantitative and qualitative experiments successfully validate the effectiveness of our proposed method. Qi Zhu 0010, Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Learning Semantic Degradation-Aware Guidance for Recognition-Driven Unsupervised Low-Light Image EnhancementabstractLow-light images suffer severe degradation of low lightness and noise corruption, causing unsatisfactory visual quality and visual recognition performance. To solve this problem while meeting the unavailability of paired datasets in wide-range scenarios, unsupervised low-light image enhancement (ULLIE) techniques have been developed. However, these methods are primarily guided to alleviate the degradation effect on visual quality rather than semantic levels, hence limiting their performance in visual recognition tasks. To this end, we propose to learn a Semantic Degradation-Aware Guidance (SDAG) that perceives the low-light degradation effect on semantic levels in a self-supervised manner, which is further utilized to guide the ULLIE methods. The proposed SDAG utilizes the low-light degradation factors as augmented signals to degrade the low-light images, and then capture their degradation effect on semantic levels. Specifically, our SDAG employs the subsequent pre-trained recognition model extractor to extract semantic representations, and then learns to self-reconstruct the enhanced low-light image and its augmented degraded images. By constraining the relative reconstruction effect between the original enhanced image and the augmented formats, our SDAG learns to be aware of the degradation effect on semantic levels in a relative comparison manner. Moreover, our SDAG is general and can be plugged into the training paradigm of the existing ULLIE methods. Extensive experiments demonstrate its effectiveness for improving the ULLIE approaches on the downstream recognition tasks while maintaining a competitive visual quality. Code will be available at https://github.com/zheng980629/SDAG. Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Zizheng Yang, Qi Zhu 0010, Feng Zhao 0004 |
AAAI | 5 |
| 2023 | Frequency-consistent Optimization for Image Enhancement Networks
Bing Li 0024, Naishan Zheng, Qi Zhu 0010, Jie Huang 0017, Feng Zhao 0004 |
BMVC | 3 |
| 2023 | Exploring Temporal Frequency Spectrum in Deep Video DeblurringabstractVideo deblurring aims to restore the latent video frames from their blurred counterparts. Despite the remarkable progress, most promising video deblurring methods only investigate the temporal priors in the spatial domain and rarely explore their its potential in the frequency domain. In this paper, we revisit the blurred sequence in the Fourier space and figure out some intrinsic frequency-temporal priors that imply the temporal blur degradation can be accessibly decoupled in the potential frequency domain. Based on these priors, we propose a novel Fourier-based frequency-temporal video deblurring solution, where the core design accommodates the temporal spectrum to a popular video deblurring pipeline of feature extraction, alignment, aggregation, and optimization. Specifically, we design a Spectrum Prior-guided Alignment module by leveraging enlarged blur information in the potential spectrum to mitigate the blur effects on the alignment. Then, Temporal Energy prior-driven Aggregation is implemented to replenish the original local features by estimating the temporal spectrum energy as the global sharpness guidance. In addition, the customized frequency loss is devised to optimize the proposed method for decent spectral distribution. Extensive experiments demonstrate that our model performs favorably against other state-of-the-art methods, thus confirming the effectiveness of frequency-temporal prior modeling. Qi Zhu 0010, Man Zhou 0003, Naishan Zheng, Chongyi Li, Jie Huang 0017, Feng Zhao 0004 |
ICCV | 1 |
| 2023 | Learning Non-Uniform-Sampling for Ultra-High-Definition Image EnhancementabstractUltra-high-definition (UHD) image enhancement is a challenging problem that aims to effectively and efficiently recover clean UHD images. To maintain efficiency, the straightforward approach is to downsample and perform most computations on low-resolution images. However, previous studies typically rely on the uniform and content-agnostic downsampling method that equally treats various regions regardless of their complexities, thus limiting the detail reconstruction in UHD image enhancement. To alleviate this issue, we propose a novel spatial-variant and invertible non-uniform downsampler that adaptively adjusts the sampling rate according to the richness of details. It magnifies important regions to preserve more information (e.g., sparse sampling points for sky, dense sampling points for buildings). Therefore, we propose a novel Non-uniform-Sampling Enhancement Network (NSEN) consisting of two core designs: 1) content-guided downsampling that extracts texture representation to guide the sampler to perform content-aware downsampling for producing detail-preserved low-resolution images; 2) invertible pixel-alignment which remaps the forward sampling process in an iterative manner to eliminate the deformations caused by the non-uniform downsampling, thus producing detail-rich clean UHD images. To demonstrate the superiority of our proposed model, we conduct extensive experiments on various UHD enhancement tasks. The results show that the proposed NSEN yields better performance against other state-of-the-art methods both visually and quantitatively. Qi Zhu 0010, Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004 |
ACM Multimedia | 2 |
| 2023 | FouriDown: Factoring Down-Sampling into Shuffling and SuperposingabstractSpatial down-sampling techniques, such as strided convolution, Gaussian, and Nearest down-sampling, are essential in deep neural networks. In this study, we revisit the working mechanism of the spatial down-sampling family and analyze the biased effects caused by the static weighting strategy employed in previous approaches. To overcome this limitation, we propose a novel down-sampling paradigm in the Fourier domain, abbreviated as FouriDown, which unifies existing down-sampling techniques. Drawing inspiration from the signal sampling theorem, we parameterize the non-parameter static weighting down-sampling operator as a learnable and context-adaptive operator within a unified Fourier function. Specifically, we organize the corresponding frequency positions of the 2D plane in a physically-closed manner within a single channel dimension. We then perform point-wise channel shuffling based on an indicator that determines whether a channel's signal frequency bin is susceptible to aliasing, ensuring the consistency of the weighting parameter learning. FouriDown, as a generic operator, comprises four key components: 2D discrete Fourier transform, context shuffling rules, Fourier weighting-adaptively superposing rules, and 2D inverse Fourier transform. These components can be easily integrated into existing image restoration networks. To demonstrate the efficacy of FouriDown, we conduct extensive experiments on image de-blurring and low-light image enhancement. The results consistently show that FouriDown can provide significant performance improvements. We will make the code publicly available to facilitate further exploration and application of FouriDown. Qi Zhu 0010, Man Zhou 0003, Jie Huang 0017, Naishan Zheng, Hongzhi Gao, Chongyi Li, Feng Zhao 0004 |
NeurIPS | 1 |
| 2022 | Dast-Net: Depth-Aware Spatio-Temporal Network for Video DeblurringabstractVideo deblurring is a challenging task due to inevitable blurs caused by depth variation, object motion, and camera shake. Although several video deblurring methods resort to depth maps, they rarely produce visually appealing results since the information in the depth maps is used insufficiently. To address this issue, we propose a Depth-Aware Modulated (DAM) block for efficiently utilizing the depth map characteristics, in which the intensity and variation of depth are exploited according to the depth map value and edges. Based on the DAM block, we develop the Depth-Aware Spatio-Temporal Network (DAST-Net) tailored for video deblurring. Particularly, the Depth-Aware Temporal Alignment module uses the depth cues to guide the alignment of adjacent frames. The Depth-Modulated Spatial Fusion module then warps the aligned frames to maintain spatial invariance with the aligned features. The warped depth features are more effective in video deblurring, since they allow for the aggregation of multiple frames. Extensive quantitative and qualitative evaluations demonstrate that the proposed DAST-Net outperforms other state-of-the-art methods. Qi Zhu 0010, Zeyu Xiao 0002, Jie Huang 0017, Feng Zhao 0004 |
ICME | 1 |
| 2022 | Source-Free Domain Adaptation for Real-World Image DehazingabstractDeep learning-based source dehazing methods trained on synthetic datasets have achieved remarkable performance but suffer from dramatic performance degradation on real hazy images due to domain shift. Although certain Domain Adaptation (DA) dehazing methods have been presented, they inevitably require access to the source dataset to reduce the gap between the source synthetic and target real domains. To address these issues, we present a novel Source-Free Unsupervised Domain Adaptation (SFUDA) image dehazing paradigm, in which only a well-trained source model and an unlabeled target real hazy dataset are available. Specifically, we devise the Domain Representation Normalization (DRN) module to make the representation of real hazy domain features match that of the synthetic domain to bridge the gaps. With our plug-and-play DRN module, unlabeled real hazy images can adapt existing well-trained source networks. Besides, the unsupervised losses are applied to guide the learning of the DRN module, which consists of frequency losses and physical prior losses. Frequency losses provide structure and style constraints, while the prior loss explores the inherent statistic property of haze-free images. Equipped with our DRN module and unsupervised loss, existing source dehazing models are able to dehaze unlabeled real hazy images. Extensive experiments on multiple baselines demonstrate the validity and superiority of our method visually and quantitatively. Hu Yu 0001, Jie Huang 0017, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004 |
ACM Multimedia | 4 |
| 2022 | Enhancement by Your Aesthetic: An Intelligible Unsupervised Personalized Enhancer for Low-Light ImagesabstractLow-light image enhancement is an inherently subjective process whose targets vary with the user's aesthetic. Motivated by this, several personalized enhancement methods have been investigated. However, the enhancement process based on user preferences in these techniques is invisible, i.e., a "black box". In this work, we propose an intelligible unsupervised personalized enhancer (iUP-Enhancer) for low-light images, which establishes the correlations between the low-light and the unpaired reference images with regard to three user-friendly attributions (brightness, chromaticity, and noise). The proposed iUP-Enhancer is trained with the guidance of these correlations and the corresponding unsupervised loss functions. Rather than a "black box" process, our iUP-Enhancer presents an intelligible enhancement process with the above attributions. Extensive experiments demonstrate that the proposed algorithm produces competitive qualitative and quantitative results while maintaining excellent flexibility and scalability. This can be validated by personalization with single/multiple references, cross-attribution references, or merely adjusting parameters. Naishan Zheng, Jie Huang 0017, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004, Zhengjun Zha |
ACM Multimedia | 3 |