EDBT 2026 Demo / reviewers in the wild / expert
Zhiwei Zhong 0001
dblp:160/7296-1
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-7716-8261ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anchor-Induced Serial Tensor Representation for Multi-View ClusteringabstractMulti-view clustering (MVC) has emerged as a powerful approach for integrating diverse sources of information from complex datasets. Nevertheless, existing methods struggle to accurately capture the global correlations and high-order structures in the data, and employ anchor-based techniques within a single dimension, limiting their representation. To address these issues, we propose an Anchor-induced Serial Tensor Representation (ASTR) framework, which effectively harnesses serial tensor representation to capture comprehensive multi-view information while reducing approximation errors and enhancing clustering performance. Specifically, ASTR begins with projection learning to explore low-dimensional latent spaces in multi-view data. Then, we introduce multi-anchor learning, where multiple anchor configurations are generated within the latent spaces, yielding a set of corresponding bipartite graphs. Besides, we organize these bipartite graphs into a sequence of global tensors, forming the serial tensor representation that encapsulates high-order inter- and intra-view relationships. Furthermore, we introduce the Laplace function to achieve a more accurate tensor rank approximation, complemented by a thorough theoretical analysis. Finally, a one-step clustering process, guided by adaptive weights, directly fuses the learned graphs to produce the final clustering indicator matrix. Experimental results demonstrate that ASTR possesses superior clustering accuracy and comparable efficiency. Zonglin Liu 0001, Zhiwei Zhong 0001, Qiangqiang Shen, Yongsheng Liang 0001, Yongyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Unfolding High-Order Correlations for Interpretable Multi-Contrast MRI Super-ResolutionabstractDeep unfolding network has gained significant attention for magnetic resonance imaging super-resolution (MRI SR) due to its performance and interpretability. However, 1) existing methods predominantly focus on cross-contrast correlations while neglecting high-order correlations embedded within spatially adjacent slices in volumetric MRI data. 2) Their degradation models are optimized via the proximal gradient algorithm (PGA) that relies on manually designed hyperparameters (e.g., step size), often leading to overshooting or suboptimal solutions. To solve these limitations, we propose HocMRI, a deep unfolding multi-contrast MRI SR framework, which seamlessly integrates dual-prior modeling and hyperparameter-free PGA for enhanced reconstruction. Specifically, we first design a novel degradation model based on the dual-prior mechanism: an explicit prior based on low-rank tensor factorization to capture intra- and inter-slice dependencies, and an implicit prior leveraging a Mamba-based network with a novel 3D scanning strategy to further exploit high-order correlations across slices. Then, we derive a hyperparameter-free PGA to boost the traditional PGA, which employs a hyperbolic tangent function to dynamically control the gradient descent step, eliminating manual tuning while ensuring stable convergence with theoretical proofs. Based on the hyperparameter-free PGA, we develop an efficient iterative optimization algorithm to solve the degradation model and unfold it into a multi-stage deep network. Numerous experimental results from widely used MRI datasets demonstrate that our HocMRI achieves superior performance with enhanced efficiency compared to the state-of-the-art methods. Qiangqiang Shen, Xuanqi Zhang, Peilin Chen 0001, Zhiwei Zhong 0001, Howard Leung, Shiqi Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration
Yudong Mao, Zhiwei Zhong 0001, Peilin Chen 0001, Zhijiang Zhang, Shiqi Wang 0001 |
CVPR | 3 |
| 2025 | Exploiting Long and Short Temporal Dependence for Low-Light Video EnhancementabstractExisting learning-based methods often lack temporal coherence in low-light video enhancement due to rarely considering intrinsic temporal dependence. To address this issue, we propose the Long-short Temporal Filtering Network (TFNet) to learn the mapping from low-light videos to normal-light ones, utilizing the well-considered data-centric strategy and a refined architecture. From the data-centric temporal strategy, we incorporate both long-range and short-range temporal dependence into TFNet, effectively capturing the temporal information. From the model design perspective, the TFNet incorporates the Temporal-aware Attentional Filtering (TAF) module, which aims to estimate and adaptively combine filtering kernels for guided filtering towards features of the middle frame. To further refine the filtered features, the cascaded Grouped Attention (GA) blocks are presented in a grouped attention strategy. Experimental results on benchmark datasets have demonstrated the superiority of our TFNet against the state-of-the-art methods in terms of video frame quality and brightness consistency. Lingyu Zhu 0006, Yudong Mao, Zhiwei Zhong 0001, Shanshe Wang, Shiqi Wang 0001 |
ICME | 5 |
| 2025 | MS-MoE: Multi-modal Structural Mixture of Experts Framework for Pan-SharpeningabstractPan-sharpening aims to generate the high-resolution (HR) multi-spectral (MS) target image from its low-resolution (LR) counterpart, which is guided by corresponding HR panchromatic (PAN) image with abundant texture structural details. Although the existing state-of-the-art methods have made remarkable progress, they are still struggling with integrating inherent structural correlation between PAN and MS images through the early or late-stage fusion alone. This would lead to texture-less pan-sharpening reconstruction due to the insufficient learning of complementary features from PAN image. To address this issue, we propose the Multi-modal Structural Mixture of Experts (MS-MoE) framework for pan-sharpening. Specifically, given the upsampled LRMS and PAN images spatially rotated at various angles, we design a set of structural experts to extract the complementary spatial and spectral features between them, in which the Texture Enhancement Module (TEM) is introduced to extract and enhance texture-structural features from different modalities. Subsequently, we introduce an additional expert network to perform feature fusion by integrating the outputs from multiple experts. To reconstruct the high-frequency information, we further leverage the Frequency feature Refinement Module (FRM) to aggregate and refine the fused features in the frequency domain. Experimental results on the benchmark pan-sharpening datasets demonstrate that the proposed MS-MoE framework achieves more competitive performance than recent state-of-the-art methods. Zhiwei Zhong 0001, Lingyu Zhu 0006, Yudong Mao, Shiqi Wang 0001 |
IJCNN | 2 |
| 2025 | Dual-Level Cross-Modality Neural Architecture Search for Guided Image Super-ResolutionabstractGuided image super-resolution (GISR) aims to reconstruct a high-resolution (HR) target image from its low-resolution (LR) counterpart with the guidance of a HR image from another modality. Existing learning-based methods typically employ symmetric two-stream networks to extract features from both the guidance and target images, and then fuse these features at either an early or late stage through manually designed modules to facilitate joint inference. Despite significant performance, these methods still face several issues: i) the symmetric architectures treat images from different modalities equally, which may overlook the inherent differences between them; ii) lower-level features contain detailed information while higher-level features capture semantic structures. However, determining which layers should be fused and which fusion operations should be selected remain unresolved; iii) most methods achieve performance gains at the cost of increased computational complexity, so balancing the trade-off between computational complexity and model performance remains a critical issue. To address these issues, we propose a Dual-level Cross-modality Neural Architecture Search (DCNAS) framework to automatically design efficient GISR models. Specifically, we propose a dual-level search space that enables the NAS algorithm to identify effective architectures and optimal fusion strategies. Moreover, we propose a supernet training strategy that employs a pairwise ranking loss trained performance predictor to guide the supernet training process. To the best of our knowledge, this is the first attempt to introduce the NAS algorithm into GISR tasks. Extensive experiments demonstrate that the discovered model family, DCNAS-Tiny and DCNAS, achieve significant improvements on several GISR tasks, including guided depth map super-resolution, guided saliency map super-resolution, guided thermal image super-resolution, and pan-sharpening. Furthermore, we analyze the architectures searched by our method and provide some new insights for future research. Zhiwei Zhong 0001, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Shiqi Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Deep Attentional Guided Image FilteringabstractGuided filter is a fundamental tool in computer vision and computer graphics, which aims to transfer structure information from the guide image to the target image. Most existing methods construct filter kernels from the guidance itself without considering the mutual dependency between the guidance and the target. However, since there typically exist significantly different edges in two images, simply transferring all structural information from the guide to the target would result in various artifacts. To cope with this problem, we propose an effective framework named deep attentional guided image filtering, the filtering process of which can fully integrate the complementary information contained in both images. Specifically, we propose an attentional kernel learning module to generate dual sets of filter kernels from the guidance and the target and then adaptively combine them by modeling the pixelwise dependency between the two images. Meanwhile, we propose a multiscale guided image filtering module to progressively generate the filtering result with the constructed kernels in a coarse-to-fine manner. Correspondingly, a multiscale fusion strategy is introduced to reuse the intermediate results in the coarse-to-fine process. Extensive experiments show that the proposed framework compares favorably with the state-of-the-art methods in a wide range of guided image filtering applications, such as guided super-resolution (SR), cross-modality restoration, and semantic segmentation. Moreover, our scheme achieved the first place in the real depth map SR challenge held in ACM ICMR 2021. The codes can be found at https://github.com/zhwzhong/DAGF. Zhiwei Zhong 0001, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Xiangyang Ji |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Spatial-Frequency Mutual Learning for Face Super-ResolutionabstractFace super-resolution (FSR) aims to reconstruct high-resolution (HR) face images from the low-resolution (LR) ones. With the advent of deep learning, the FSR technique has achieved significant breakthroughs. However, existing FSR methods either have a fixed receptive field or fail to maintain facial structure, limiting the FSRperformance. To circumvent this problem, Fourier transform is introduced, which can capture global facial structure information and achieve image-size receptive field. Relying on the Fourier transform, we devise a spatial-frequency mutual network (SFMNet) for FSR, which is the first FSR method to explore the correlations between spatial and frequency domains as far as we know. To be specific, our SFMNet is a two-branch network equipped with a spatial branch and a frequency branch. Benefiting from the property of Fourier transform, the frequency branch can achieve image-size receptive field and capture global dependency while the spatial branch can extract local dependency. Considering that these dependencies are complementary and both favorable for FSR, we further develop a frequency-spatial interaction block (FSIB) which mutually amalgamates the complementary spatial and frequency information to enhance the capability of the model. Quantitative and qualitative experimental results show that the proposed method out-performs state-of-the-art FSR methods in recovering face images. The implementation and model will be released at https://github.com/wcy-cs/SFMNet. Chenyang Wang 0002, Junjun Jiang, Zhiwei Zhong 0001, Xianming Liu 0005 |
CVPR | 3 |
| 2022 | Self-Supervised Arbitrary-Scale Point Clouds Upsampling via Implicit Neural RepresentationabstractPoint clouds upsampling is a challenging issue to gener-ate dense and uniform point clouds from the given sparse input. Most existing methods either take the end-to-end su-pervised learning based manner, where large amounts of pairs of sparse input and dense ground-truth are exploited as supervision information; or treat up-scaling of different scale factors as independent tasks, and have to build multiple networks to handle upsampling with varying factors. In this paper, we propose a novel approach that achieves self-supervised and magnification-flexible point clouds upsampling simultaneously. We formulate point clouds upsampling as the task of seeking nearest projection points on the implicit surface for seed points. To this end, we define two implicit neural functions to estimate projection direction and distance respectively, which can be trained by two pretext learning tasks. Experimental results demonstrate that our self-supervised learning based scheme achieves competitive or even better performance than supervised learning based state-of-the-art methods. The source code is publicly available at https://github.com/xnowbzhaolsapcu. Wenbo Zhao 0004, Xianming Liu 0005, Zhiwei Zhong 0001, Junjun Jiang, Wei Gao 0003, Ge Li 0002, Xiangyang Ji |
CVPR | 3 |
| 2022 | Propagating Facial Prior Knowledge for Multitask Learning in Face Super-ResolutionabstractExisting face hallucination methods always achieve improved performance through regularizing the model with facial prior. Most of them always estimate facial prior information first and then leverage it to help the prediction of the target high-resolution face image. However, the accuracy of prior estimation is difficult to guarantee, especially for the low-resolution face image. Once the estimated prior is inaccurate or wrong, the following face super-resolution performance is unavoidably influenced. A natural question that arises: how to incorporate facial prior effectively and efficiently without prior estimation? To achieve this goal, we propose to learn facial prior knowledge at training stage, but test only with low-resolution face image, which can overcome the difficulty of estimating accurate prior. In addition, instead of estimating facial prior, we directly explore the potential of high-quality facial prior in the training phase and progressively propagate the facial prior knowledge from the teacher network (trained with the low-resolution face/high-quality facial prior and high-resolution face image pairs) to the student network (trained with the low-resolution face and high-resolution face image pairs). Quantitative and qualitative comparisons on benchmark face datasets demonstrate that our method outperforms the state-of-the-art face super-resolution methods. The source codes of the proposed method will be available athttps://github.com/wcy-cs/KDFSRNet. Chenyang Wang 0002, Junjun Jiang, Zhiwei Zhong 0001, Xianming Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | High-Resolution Depth Maps Imaging via Attention-Based Hierarchical Multi-Modal FusionabstractDepth map records distance between the viewpoint and objects in the scene, which plays a critical role in many real-world applications. However, depth map captured by consumer-grade RGB-D cameras suffers from low spatial resolution. Guided depth map super-resolution (DSR) is a popular approach to address this problem, which attempts to restore a high-resolution (HR) depth map from the input low-resolution (LR) depth and its coupled HR RGB image that serves as the guidance. The most challenging issue for guided DSR is how to correctly select consistent structures and propagate them, and properly handle inconsistent ones. In this paper, we propose a novel attention-based hierarchical multi-modal fusion (AHMF) network for guided DSR. Specifically, to effectively extract and combine relevant information from LR depth and HR guidance, we propose a multi-modal attention based fusion (MMAF) strategy for hierarchical convolutional layers, including a feature enhancement block to select valuable features and a feature recalibration block to unify the similarity metrics of modalities with different appearance characteristics. Furthermore, we propose a bi-directional hierarchical feature collaboration (BHFC) module to fully leverage low-level spatial information and high-level structure information among multi-scale features. Experimental results show that our approach outperforms state-of-the-art methods in terms of reconstruction accuracy, running speed and memory efficiency. Zhiwei Zhong 0001, Xianming Liu 0005, Junjun Jiang, Debin Zhao, Zhiwen Chen 0002, Xiangyang Ji |
IEEE Trans. Image Process. | 1 |
| 2020 | Parsing Map Guided Multi-Scale Attention Network For Face HallucinationabstractFace hallucination that aims to transform a low-resolution (LR) face image to a high-resolution (HR) one is an active domain-specific image super-resolution problem. The performance of existing methods is usually not satisfactory, especially when the upscaling factor is large, such as 8×. In this paper, we propose an effective two- step face hallucination method based on a deep neural network with multi-scale channel and spatial attention mechanism. Specifically, we develop a ParsingNet to extract the prior knowledge of an input LR face, which is then fed into a carefully designed FishSRNet to recover the target HR face. Experimental results demonstrate that our method outperforms the state-of-the-arts in terms of quantitative metrics and visual quality. Chenyang Wang 0002, Zhiwei Zhong 0001, Junjun Jiang, Deming Zhai, Xianming Liu 0005 |
ICASSP | 2 |