EDBT 2026 Demo / reviewers in the wild / expert
Lingchen Gu
dblp:205/3335
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-3127-3119ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCIA-GAN: Robust image watermarking via spatial-channel interaction attention and feature preservation
Lingchen Gu, Jun Wang 0061, Wenbo Wan, Jiande Sun 0001, Sen-Ching S. Cheung |
Expert Syst. Appl. | 2 |
| 2026 | DNA-Former: Video DNA-Aware Transformer for Human Action RecognitionabstractABSTRACT Human action recognition is achieved by extracting category‐based common semantic features in videos. However, most of the current transformers draw attention to the space‐time relationship of visual patches, with limited ability to perceive semantic connections between them. We notice that the visual expression of videos comes from semantics, which is similar to the expression mechanism of human deoxyribonucleic acid (DNA). To mitigate the considerable gap between video visual features and semantic information, we build a sparse video DNA space by mimicking the human DNA, and present a video transformer termed as DNA‐Former. Specially, we design a DNA dual‐attention block to learn the correlation between video DNA information and visual patches in the space‐time dimension. In this block, we explore the cross‐shaped window self‐attention, which is rarely applied in video transformers. To further select more expressive visual patch features, we develop the contrastive disentangler head to filter out the DNA‐aware useless features based on contrastive learning. Our proposed DNA‐Former is performed on different benchmarks (i.e. Kinetics‐400, Something‐Something V2, UCF‐101 and HMDB‐51) with comprehensive ablation studies and visualization analysis. Experimental results demonstrate the effectiveness of our method, showing favourable accuracy against state‐of‐the‐art methods. Aixi Qu, Lingchen Gu |
IET Image Process. | 3 |
| 2026 | A universal pansharpening network via spatial-spectral contrastive learning
Kai Zhang 0010, Yunlong Liu 0005, Feng Zhang 0028, Wenbo Wan, Lingchen Gu, Jiande Sun 0001 |
Pattern Recognit. | 6 |
| 2026 | Quality-Guided Forgery Adapter for Generalizable AIGC Image DetectionabstractThe rapid advancement of AI-generated content (AIGC) presents significant challenges for digital forensics, necessitating robust and generalizable detection frameworks. Existing detection methods primarily rely on visual feature extraction, while vision-language model-based approaches are limited to class-label prompts, failing to capture quality-related artifacts introduced by different generative models. To address this limitation, we introduce QAFD, a novel Quality-Assisted Forgery Detection framework that incorporates image quality information into the detection process. Specifically, we design a quality queried attention block to effectively fuse class-based content prompts with quality-aware text prompts. This integration enhances the model’s ability to capture semantic artifacts related to degradation patterns commonly associated with AI-generated images. Furthermore, we introduce the Quality-Guided Forgery Adapter (QGFA) to incorporate quality-aware textual cues into the visual domain, improving feature extraction for both spatial and frequency-based forgery artifacts. This synergy allows frequency cues to enhance low-level artifact perception, while quality-aware guidance strengthens high-level discriminative representation. Extensive experiments demonstrate that QAFD achieves superior generalization to unseen generative models over three datasets and significantly maintains its robustness against common image post-processing operations.The codes will be released at github. Jun Wang 0061, Zitong Yu, Chaomeng Chen, Lingchen Gu, Wenbo Wan, Jiantao Zhou 0001, Weiming Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | WPM-GAN: Watermark-Preserving Module for GAN-Based Robust Industrial Image WatermarkingabstractIn this paper, a robust watermarking framework is proposed for ensuring copyright and authenticity of industrial imagery. Leveraging Squeeze-and-Excitation (SE) blockbased encoder-decoder network, our method embeds and extracts watermarks in a more efficient and imperceptible manner, and our method introduces a novel Watermark-Preserving Module (WPM) at the receiving end, which maximizes the preservation of watermark features in noise-corrupted images to assist the decoder in watermark extraction. And we adopt a dualstage training strategy to capture and learn specific watermark features, along with a decoder-based watermark-preserving loss to further enhance robustness. Experimental results demonstrate that WPM-GAN achieves superior visual quality while effectively resisting various attacks. Yingchao Yang, Lingchen Gu, Wenbo Wan |
ICPADS | 4 |
| 2025 | Mining the Salient Spatio-Temporal Feature with S2TF-Net for action recognition
Lingchen Gu, Xiaojun Chang, Feiping Nie 0001 |
Signal Process. Image Commun. | 3 |
| 2025 | Dual Prototypes-Based Personalized Federated Adversarial Cross-Modal HashingabstractWith the rapid advances in wireless communication and IoT platforms, it is increasingly difficult to analyze relevant multi-modal data distributed across geographically diverse and heterogeneous platforms. One promising approach is to rely on federated learning to build compact cross-modal hash codes. However, existing federated learning methods easily exhibit degenerative performance in the global model due to the distributed data being derived from diverse domains. In addition, directly forcing each client to adopt the same global parameters as local parameters, without effective local training, significantly reduces the performance of each client. To overcome these challenges, we propose a novel federated adversarial cross-modal hashing, called Dual Prototypes-based personalized Federated Adversarial (DP-FeAd), which provides iterated training of shared dual prototypes. Specifically, aiming to expand local hashing models beyond their knowledge realms, DP-FeAd enables participating clients to engage in cooperative learning through two constructions: cluster prototypes and unbiased prototypes, instead of the traditional global prototypes, ensuring both generalization and stability. Specifically, the cluster prototypes are derived from local class-level prototypes and adversarially trained with local approximate hash codes to align their distributions. The unbiased prototypes are averaged from cluster prototypes and integrated into the training of local hashing models to maintain consistency across different local class-level prototypes further. The experiments conducted on two benchmark datasets demonstrate that our proposed method significantly enhances the performance of deep cross-modal hashing models in both IID (Independent and Identically Distributed) and non-IID scenarios. Lingchen Gu, Xiaojuan Shen, Jiande Sun 0001, Jing Li 0046, Zhihui Li 0001, Sen-Ching S. Cheung, Wenbo Wan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Entropy-Optimized Deep Weighted Product Quantization for Image RetrievalabstractHashing and quantization have greatly succeeded by benefiting from deep learning for large-scale image retrieval. Recently, deep product quantization methods have attracted wide attention. However, representation capability of codewords needs to be further improved. Moreover, since the number of codewords in the codebook depends on experience, representation capability of codewords is usually imbalanced, which leads to redundancy or insufficiency of codewords and reduces retrieval performance. Therefore, in this paper, we propose a novel deep product quantization method, named Entropy Optimized deep Weighted Product Quantization (EOWPQ), which not only encodes samples into the weighted codewords in a new flexible manner but also balances the codeword assignment, improving while balancing representation capability of codewords. Specifically, we encode samples using the linear weighted sum of codewords instead of a single codeword as traditionally. Meanwhile, we establish the linear relationship between the weighted codewords and semantic labels, which effectively maintains semantic information of codewords. Moreover, in order to balance the codeword assignment, that is, avoiding some codewords representing most samples or some codewords representing very few samples, we maximize the entropy of the coding probability distribution and obtain the optimal coding probability distribution of samples by utilizing optimal transport theory, which achieves the optimal assignment of codewords and balances representation capability of codewords. The experimental results on three benchmark datasets show that EOWPQ can achieve better retrieval performance and also show the improvement of representation capability of codewords and the balance of codeword assignment. Lingchen Gu, Wenbo Wan, Jiande Sun 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Video Summarization Through Fine-Grained Hierarchical Modeling with Multi-Dimensional FeaturesabstractVideo summarization aims to shorten the video length while maintaining the original video content, which facilitates large-scale video searching and browsing. Most of the existing methods simply take static image features as input, which causes the loss of temporal action information of successive frames. Additionally, the use of two-stage temporal modeling aggravates the loss of temporal relationship. In this paper, we propose a framework based on Fine-Grained Hierarchical Modeling (FGHM) employing multi-dimensional features. Firstly, the multi-dimensional features extractor extracts static image features and dynamic video features. Then dynamic temporal modeling is carried out to model the temporal dependency of the entire video. We also investigate the effects of spatial-temporal features extracted by various 3D features extractors. Extensive experiments demonstrate the effectiveness of FGHM against state-of-the-art methods. Mengnan Liang, Lingchen Gu |
ICIP | 4 |
| 2023 | AFcIHNet: Attention feature-constrained network for single image information hiding
Xingwang Jia, Hua-Mei Xin 0001, Lingchen Gu, Jiande Sun 0001, Wenbo Wan |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | Enhancing adversarial transferability with partial blocks on vision transformer
Yanyang Han, Lingchen Gu, Xuesong Gao |
Neural Comput. Appl. | 5 |
| 2022 | Dual Distance Optimized Deep Quantization With Semantics-PreservingabstractRecently, quantization has been an effective technique for large-scale image retrieval, which can encode feature vectors into compact codes. However, it is still a great challenge to improve the discriminative capability of codewords while minimizing the quantization error. This letter proposes Dual Distance Optimized Deep Quantization (D2ODQ) to deal with this issue, by minimizing the Euclidean distance between samples and codewords, and maximizing the minimum cosine distance between codewords. To generate the evenly distributed codebook, we find the general solution for the upper bound of the minimum cosine distance between codewords. Moreover, scaler constrained semantics-preserving loss is considered to avoid trivial quantization boundary, and ensure that a codeword can only quantize the features of one category. In contrast to state-of-the-art methods, our method has a better performance on three benchmark datasets. Lingchen Gu, Zhengfeng Du |
IEEE Signal Process. Lett. | 1 |
| 2021 | Deep image hashing based on twin-bottleneck hashing with variational autoencodersabstractWith the ever-increasing availability of data, the need for efficient and accurate image retrieval methods has become larger and larger. Deep hashing has proven to be a promising solution, by defining a hash function to convert the data into a manageable lower-dimensional representation. In this paper, we apply recent insights from the field of variational autoencoders to the field of deep image hashing, thus achieving an improvement over the current state of the art as shown by experimental evaluation. The code used in this paper is open-source and available on GitHub (https://github.com/maximverwilst/deepimagehashing-VAE). Maxim Verwilst, Nina Zizakic, Lingchen Gu, Aleksandra Pizurica |
MMSP | 3 |
| 2021 | Deep Loss Driven Multi-Scale Hashing Based on Pyramid Connected NetworkabstractThanks to the great success of the deep learning, deep hashing for large-scale multimedia retrieval has made significant progress recently. However, most existing deep hashing algorithms suffer from slow convergence due to the gradient vanishing problem, caused by deep network structures and saturated activation functions. Moreover, a single convolution layer is often followed by down-sampling such as max pooling, resulting in local information loss that might affect the overall system robustness and performance. In this work, we propose a novel deep supervised hashing, Deep Loss Driven Multi-Scale Hashing (DLDMSH), which learns the high-quality approximate binary codes through an end-to-end network and improves the representative capacity of hash codes for large-scale image retrieval. Specifically, we design a Loss Driven Multi-Scale (LDMS) feature which is aggregated from convolutional feature maps. Moreover, a Pyramid Connected Convolutional Neural Network (PCNet) architecture is devised to generate LDMS feature, which inputs pairs of images during the training and outputs an image to approximate discrete values. In particular, 1 × 1 convolution kernels are applied to make a linear combination of features for realizing feature reduction, and the reduced features are fused in the fusion layer. This effectively improves the performance of deep features. A novel loss function preserving semantic information is integrated into an end-to-end learning scheme, which enhances the representative capacity of binary codes. Extensive experiments over four benchmark datasets show that DLDMSH significantly outperforms several other state-of-the-art hashing methods. Lingchen Gu, Jiande Sun 0001 |
IEEE Trans. Multim. | 1 |