VLDB 2026 Research / reviewers in the wild / expert
Liguo Zhang 0002
dblp:52/952-2
· DBLP profile ↗
28ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-3814-7783ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Computer networks · 7 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Noise Modeling for Spike Camera via Time-Interval Quantification and Spike-DSLR Multimodal Dataset in Low-Light ImagingabstractThe inherent differences between spike cameras and traditional frame-based cameras lead to more complex and diverse noise characteristics, particularly under extremely low-light conditions. Existing noise modeling approaches for spike camera predominantly rely on inter-spike intervals (ISI) for noise quantification, which often results in inaccurate noise characterization. Moreover, current datasets for spike camera image reconstruction tasks are either synthetic or lack corresponding high-quality reference images, severely limiting rigorous evaluation of noise modeling methods. To address this limitation, we propose a multimodal noise modeling framework for spike camera that integrates insights from traditional frame-based imaging into spike imaging. Specifically, we introduce a time-interval-based quantification method inspired by the exposure-time concept used in traditional frame-based cameras, enabling accurate noise characterization for spike camera. Furthermore, we present the Spike-DSLR Multimodal Dataset (SDMD), the first real-world dataset capturing aligned multimodal data pairs from spike cameras and Digital Single-Lens Reflex (DSLR) cameras, explicitly designed for evaluating spike camera noise models. Experimental results on SDMD demonstrate that our noise modeling approach significantly enhances spike camera image reconstruction quality under low-light conditions, achieving more than 1.6 dB improvement in PSNR compared to existing state-of-the-art methods. This validates both the necessity and effectiveness of adopting a multimodal perspective in spike camera noise modeling. Yue Cao 0009, Sizhao Li, Liguo Zhang 0002 |
AAAI | 3 |
| 2026 | An enriched Transformer powered by knowledge graph for multi-task Ethereum fraud detection
Ye Tian 0027, Liguo Zhang 0002, Zhiquan Liu 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Adaptive spectral bandpass multi-scale network for underwater acoustic target recognition
Pengyuan Qi, Guisheng Yin, Yuxin Dong 0001, Liguo Zhang 0002 |
Multim. Syst. | 5 |
| 2025 | Zero-Shot Noise2Mean: Gap Minimization for Efficient Denoising from a Single Noisy ImageabstractAcquiring pairwise noisy-clean training data is challenging. Consequently, some self-supervised denoising methods utilize noisy image pairs as both input and target for network training. However, a major issue with these methods is the gap between the clean images of the input and target. In this paper, we achieve high-quality image denoising by reducing or even eliminating this gap. Our method requires no training data or prior knowledge of the noise distribution. It consists of two lightweight networks that can be trained using only a single noisy test image. Specifically, we propose a random mask-based downsampler that generates multiple pairs of downsampled noisy images, which are similar but distinct. These image pairs serve as the input for the first network, with the mean image of each pair used as the target. This initially reduces the gap between the clean images of the input and target. Particularly, in our method, the clean counterpart of the first network's target (i.e., the mean image) can be obtained. We then train a second network using the mean image as input and its clean counterpart as the target. This effectively eliminates the gap and achieves better denoising results. Extensive experiments demonstrate that our method outperforms in both denoising performance and efficiency. Yiqi Shi, Guoyin Zhang, Sizhao Li, Liguo Zhang 0002 |
AAAI | 5 |
| 2025 | CG2AN: Conditional Graph Generative Adversarial Networks for Shredded Image RestorationabstractThe automatic reconstruction of image and document from their fragments is a classic and challenging problem in image processing. It has very important theoretical significance and application value in criminal investigation, archive restoration and more. This paper introduces a novel method for the reconstruction of shredded images by combining image processing techniques with graph reconstruction strategies. First, the stitching of fragments is treated as a graph reconstruction problem. Each chip of the image is considered a vertex within a graph. Second, the deformable convolutional networks are employed to extract features in the boundary area of the fragment. Third, we propose a conditional graph generative adversarial network to generate the adjacency matrix and reconstruct the graph topology. In this process, the gradient-weighted class activation mapping is applied to capture and visualize the adjacent boundaries of fragments. By calculating the center of mass and gradient direction within the activated area, we can accurately align fragments. Finally, the matched fragments are iteratively stitched together on the basis of the reconstructed graph and adjacent boundaries. The experimental results show that compared with other algorithms, the proposed method exhibits higher precision and reliability across multiple evaluation indexes. Liguo Zhang 0002, Yuxin Dong 0001, Yalong Zhu |
ECAI | 1 |
| 2025 | ZVEFusion: Zero-Shot Visual Enhancement Fusion for Infrared and Visible Images in Low LightabstractInfrared and visible image fusion (IVIF) aims to generate fused images with prominent targets and rich scene information. However, in low-light conditions, visible images lose accurate texture and color, reducing their ability to provide detailed scene information for fusion. Existing IVIF methods often overlook illumination degradation and cause color distortion when incorporating infrared information. To address these problems, we propose a novel visually enhanced IVIF method tailored for low-light environments. Our method combines low-light image enhancement (LLIE) and IVIF into a single module. First, we adaptively enhance the low-light visible image, ensuring rich texture and color for fusion. Additionally, we introduce a three-channel fusion coefficient map to transform infrared information into visible image, preventing color distortion and highlighting key targets while maintaining details of the fused image. Since infrared and visible images are from different modalities, we map them into the same high-dimensional feature space. We then propose the feature difference to integrate complementary information, producing a fused image with complete content and no redundancy. Notably, our method is zero-shot, requiring only a pair of test infrared and visible images for training. This better meets the complexity of IVIF in various low-light scenes. Extensive experiments show that in low-light conditions, our method surpasses other state-of-the-art (SOTA) methods by providing more natural colors, richer textures, and better alignment with human visual perception. Yiqi Shi, Guoyin Zhang, Sizhao Li, Liguo Zhang 0002 |
ICASSP | 5 |
| 2025 | Boosting Adversarial Transferability via Negative Hessian Trace Regularization
Zilin Tian, Liguo Zhang 0002, Huosheng Xu |
ICCV | 3 |
| 2025 | ERTFNet: Enhanced RGB-T Fusion Network for semantic segmentation by integrating thermal edge features
Hanqi Yin, Liguo Zhang 0002, Guisheng Yin |
Comput. Vis. Image Underst. | 2 |
| 2024 | ZERO-IG: Zero-Shot Illumination-Guided Joint Denoising and Adaptive Enhancement for Low-Light ImagesabstractThis paper presents a novel zero-shot method for jointly denoising and enhancing real-word low-light images. The proposed method is independent of training data and noise distribution. Guided by illumination, we integrate denoising and enhancing processes seamlessly, enabling end-to-end training. Pairs of downsampled images are extracted from a single original low-light image and processed to preliminarily reduce noise. Based on the smoothness of illumination, near-authentic illumination can be estimated from the denoised low-light image. Specifically, the illumi-nation is constrained by the denoised image's brightness, uniformly amplifying pixels to raise overall brightness to normal-light level. We simultaneously restrict the illumi-nation by scaling each pixel of the denoised image based on its intensity, controlling the enhancement amplitude for different pixels. Applying the illumination to the original low-light image yields an adaptively enhanced reflection. This prevents under-enhancement and localized overexpo-sure. Notably, we concatenate the reflection with the illumi-nation, preserving their computational relationship, to ul-timately remove noise from the original low-light image in the form of reflection. This provides sufficient image infor-mation for the denoising procedure without changing the noise characteristics. Extensive experiments demonstrate that our method outperforms other state-of-the-art meth-ods. The source code is available at https://github.com/Doyle59217/ZeroIG. Yiqi Shi, Liguo Zhang 0002, Ye Tian 0027, Xuezhi Xia, Xiaojing Fu |
CVPR | 3 |
| 2024 | DP-Font: Chinese Calligraphy Font Generation Using Diffusion Model and Physical Information Neural Network
Liguo Zhang 0002, Yalong Zhu, Achref Benarab, Yusen Ma, Yuxin Dong 0001 |
IJCAI | 1 |
| 2024 | Time-Frequency Domain Fusion Enhancement for Audio Super-ResolutionabstractAudio super-resolution aims to improve the quality of acoustic signals and is able to reconstruct corresponding high-resolution acoustic signals from low-resolution acoustic signals. However, since acoustic signals can be divided into two forms: time-domain acoustic waves or frequency-domain spectrograms, most existing research focuses on data enhancement in a single field, which can only obtain partial or local features of the audio signal, resulting in limitations of data analysis. Therefore, this paper proposes a time-frequency domain fusion enhanced audio super-resolution method to mine the complementarity of the two representations of acoustic signals. Specifically, we propose an end-to-end audio super-resolution network. Including the variational autoencoder based sound wave super-resolution module, U-Net-based Spectrogram Super-Resolution Module, and attention-based Time-Frequency Domain Fusion Module. The first two modules can generate more high-frequency and low-frequency components for audio respectively. As a critical component of our method, time-frequency domain fusion module performs weighted fusion on the above two outputs to obtain a super-resolution audio signal. Compared with other methods, experimental results on the VCTK and Piano datasets in natural scenes show that the time-frequency domain fusion audio super-resolution model has a state-of-the-art bandwidth expansion effect. Furthermore, we perform super-resolution on the ShipsEar dataset containing underwater acoustic signals. The super-resolution results are used to test ship target recognition, and and the accuracy is improved by 12.66%. Therefore, the proposed super-resolution method has excellent signal enhancement effect and generalization ability. Ye Tian 0027, Liguo Zhang 0002 |
ACM Multimedia | 4 |
| 2024 | Adversarial Sample Optimization in Multi-DomainsabstractAdversarial attacks have attracted much attention in the field of vision. In recent years, many scholars have proposed their own views in the field of adversarial attacks. Traditional methods generate adversarial samples with very small differences in the spatial domain compared to the original samples, but in fact, there is still a large difference when transferring the adversarial samples to the frequency domain. In order to improve the ability of adversarial attacks and reduce the amount of frequency domain disturbance, this paper proposes a cross-domain optimization algorithm for adversarial samples. The main idea is to crop the perturbation range of the adversarial samples generated by adversarial attacks and the original images in the frequency domain space, and combine the optimized samples after cropping with the original samples to form new adversarial samples. Experiments show that this method eliminates the differences generated by the adversarial samples in the frequency domain space while hardly reducing the aggressiveness of the adversarial samples. It further improves the stealth of the adversarial samples with almost no burden. Yalong Zhu, Liguo Zhang 0002 |
MSN | 3 |
| 2024 | Parallel Correlation Attention Modules are Used for Feature Extraction and Fusion to Achieve Accurate Target SegmentationabstractThe conventional convolutional neural network (CNN) model has limitations in terms of global modeling ability, as it can only extract local features and is susceptible to noise interference. Due to repeated downsampling, small target features are easily lost in the deeper layers of the network. On the other hand, Transformer models are renowned for their exceptional global modeling capabilities; however, for the specific task of polyp images, effective results necessitate local feature extraction. Therefore, when reconsidering the relationship between local and global aspects within CNN and self-attention models, we propose a parallel relevant attention module with a Transformer structure for feature extraction and fusion. By combining channel attention and spatial attention mechanisms, we obtain semantic information from the lowest-level features to determine target feature locations accurately. Additionally, our step-by-step feature fusion module integrates shallow feature information into deeper layers through a CNN structure to capture more detailed target features comprehensively. Finally, these detailed target features are combined with underlying semantic information to achieve precise target segmentation. Mingze Xia, Liguo Zhang 0002, Guisheng Yin, Yuxin Dong 0001 |
MSN | 2 |
| 2024 | Multimodal Data Generation for Audio and Video InteractionabstractMultimodal Large Language Models (MM-LLMs) have made significant progress in the field of artificial intelligence, which not only inherit the advantages of traditional Large Language Models (LLMs) in text processing, but also realize the ability of cross-modal understanding and generation. This paper, based on the existing framework of MM-LLMs, realizes a cross-modal understanding and generation of LLM on audio and video. The model has four layers: modal encoder layer, input projection layer, LLM layer and modal generator layer. This model improves the input projection layer of the general MM-LLM architecture and the alignment between the LLM and the modal generator, and fine-tuning the modal generator to make the architecture more suitable for the task requirements of the model. Yalong Zhu, Liguo Zhang 0002 |
MSN | 3 |
| 2024 | Underwater acoustic target recognition using RCRNN and wavelet-auditory feature
Pengyuan Qi, Guisheng Yin, Liguo Zhang 0002 |
Multim. Tools Appl. | 3 |
| 2024 | Robust semantic segmentation method of urban scenes in snowy environment
Hanqi Yin, Guisheng Yin, Liguo Zhang 0002, Ye Tian 0027 |
Mach. Vis. Appl. | 4 |
| 2023 | Cross-modal and Cross-medium Adversarial Attack for AudioabstractAcoustic waves are forms of energy that propagate through various mediums. They can be represented by different modalities, such as auditory signals and visual patterns. The two modalities are often described as one-dimensional waveform in the time domain and two-dimensional spectrogram in the frequency domain. Most acoustic signal processing methods use single modal data for input and training models. This poses a challenge for black-box adversarial attacks on audio signals because the input modality is also unknown to the attacker. In fact, there currently exist no methods that explore the cross-modal transferability of adversarial perturbation. This paper investigates the cross-modal transferability from waveform to spectrogram. We argue that the data distributions in the sample space with the different modalities have mapping relations and propose a novel decision-based cross-modal and cross-medium adversarial attack method. Specifically, it generates an initial example with cross-modal attack capability by combining random natural noise, then iteratively reduces the perturbation to enhance its invisibility. It incorporates the constraints of the spectrogram sample space while iteratively optimizing adversarial perturbations for black-box audio classification models. The perturbation is imperceptible to humans, both visually and aurally. Extensive experiments demonstrate that our approach can launch attacks on classification models for sound waves and spectrograms that share the same audio signal. Furthermore, we explore the cross-medium capability of our proposed adversarial attack strategy that can target processing models for acoustic signals propagating in air and seawater. The proposed method has preeminent invisibility and generalization compared to other methods. Liguo Zhang 0002, Zilin Tian, Sizhao Li, Guisheng Yin |
ACM Multimedia | 1 |
| 2023 | PTKE: Translation-based temporal knowledge graph embedding in polar coordinate system
Ruinan Liu, Guisheng Yin, Zechao Liu, Liguo Zhang 0002 |
Neurocomputing | 4 |
| 2022 | Consistency regularization teacher-student semi-supervised learning method for target recognition in SAR images
Ye Tian 0027, Liguo Zhang 0002, Guisheng Yin, Yuxin Dong 0001 |
Vis. Comput. | 2 |
| 2021 | Cloud-Based Data Offloading for Multi-focus and Multi-views Image Fusion in Mobile Applications
Yiqi Shi, Liang Kou, Boquan Li 0002, Qing Yang 0003, Liguo Zhang 0002 |
Mob. Networks Appl. | 7 |
| 2021 | Deeper super-resolution generative adversarial network with gradient penalty for sonar image enhancement
Pengyang Shen, Liguo Zhang 0002, Guisheng Yin |
Multim. Tools Appl. | 2 |
| 2021 | EmotionCues: Emotion-Oriented Visual Summarization of Classroom VideosabstractAnalyzing students' emotions from classroom videos can help both teachers and parents quickly know the engagement of students in class. The availability of high-definition cameras creates opportunities to record class scenes. However, watching videos is time-consuming, and it is challenging to gain a quick overview of the emotion distribution and find abnormal emotions. In this article, we propose EmotionCues, a visual analytics system to easily analyze classroom videos from the perspective of emotion summary and detailed analysis, which integrates emotion recognition algorithms with visualizations. It consists of three coordinated views: a summary view depicting the overall emotions and their dynamic evolution, a character view presenting the detailed emotion status of an individual, and a video view enhancing the video analysis with further details. Considering the possible inaccuracy of emotion recognition, we also explore several factors affecting the emotion analysis, such as face size and occlusion. They provide hints for inferring the possible inaccuracy and the corresponding reasons. Two use cases and interviews with end users and domain experts are conducted to show that the proposed system could be useful and effective for analyzing emotions in the classroom videos. Haipeng Zeng, Xinhuan Shu, Yanbang Wang, Yong Wang 0021, Liguo Zhang 0002, Ting-Chuen Pong, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | $\hbox {S}^2\hbox {RGAN}$: sonar-image super-resolution based on generative adversarial network
Liguo Zhang 0002, Yang Li 0122, Guisheng Yin |
Vis. Comput. | 3 |
| 2020 | A data authentication scheme for UAV ad hoc network communication
Liang Kou, Yun Lin 0005, Liguo Zhang 0002, Qingan Da, Lei Chen 0029 |
J. Supercomput. | 5 |
| 2019 | A multi-focus image fusion algorithm in 5G communications
Kejia Zhang 0001, Liguo Zhang 0002, Yun Lin 0005, Qilong Han, Qingan Da, Liang Kou |
Multim. Tools Appl. | 4 |
| 2018 | A Novel Hybrid Information Security Scheme for 2D Vector Map
Qingan Da, Liguo Zhang 0002, Liang Kou, Qilong Han, Ruolin Zhou |
Mob. Networks Appl. | 3 |
| 2016 | A support vector machine based naive Bayes algorithm for spam filteringabstractNaive Bayes classifiers are widely used to filter spam emails, however, the strong independence assumptions between features limit their performance in accurately identifying spams. To address this issue, we proposed a support machine vector based naive Bayes - SVM-NB - filtering system. The SVM-NB first constructs an optimal separating hyperplane that divides samples in the training set into two categories. For samples located nearby the hyperplane, if they are in different categories, one of them will be eliminated from the training set. In this way, the dependence between samples is reduced and the entire training sample space is simplified. With the trimmed training set, the naive Bayes algorithm is applied to classify emails in the test set. The SVM-NB system is evaluated with the dataset obtained from DATAMALL. Experiment results demonstrate that SVM-NB can achieve a higher spam-detection accuracy and a faster classification speed. Weimiao Feng, Liguo Zhang 0002, Cuiling Cao, Qing Yang 0003 |
IPCCC | 3 |
| 2016 | Multi-focus Image Fusion via Region Mosaicing on Contrast Pyramids
Liguo Zhang 0002, Weimiao Feng, Qing Yang 0003 |
WASA | 1 |