VLDB 2026 Research / reviewers in the wild / expert
Guixuan Zhang
dblp:166/6447
· DBLP profile ↗
19ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-1072-8279ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Branch Asymmetric Discrepancy Learning Based on Fake Image Pattern-Coexistence for AI-Generated Image DetectionabstractWith the rapid advancement of generative models, high-fidelity AI-generated images have become increasingly indistinguishable from real images, posing significant challenges to traditional detection methods that rely on explicit artifacts or uniform feature learning. We hypothesize that detection ambiguity originates from pattern coexistence: synthetic images simultaneously embed (a) authentic patterns inherited from real-image distributions and (b) synthetic patterns induced by generative architectures, whereas real images maintain consistent patterns. We validate this hypothesis through SHAP-based quantitative analysis, demonstrating that synthetic images inherently exhibit a dual distribution—simultaneously containing authentic patterns and synthetic traces—while real images show a unimodal distribution. Building on this insight, this paper proposes a Dual-Branch Asymmetric Discrepancy Learning (DADL) framework. The DADL leverages multi-scale feature extraction and Asymmetric Feature Discrepancy Loss to capture and amplify such pattern differences across multiple scales. Extensive experiments on three benchmarks (AIGCDetectBenchmark, GenImage, and Chameleon) show that DADL achieves state-of-the-art performance, with particular strengths in detecting high-fidelity synthetic images from diffusion models (e.g., Midjourney, SDv1.4, SDv1.5) and enhancing generalization across diverse generative paradigms. This study not only offers an effective approach for AIGI detection but also sheds light on the intrinsic properties of synthetic images, providing a new perspective for advancing AIGI forensics. Chunli Song, Peiyang Wang, Guixuan Zhang, Shuwu Zhang |
AAAI | 5 |
| 2026 | Keypoint-enhanced image watermarking with spatial-frequency mapping and perceptual optimization
Fei Ge, Jie Liu 0028, Guixuan Zhang, Shuwu Zhang, Hu Guan |
Inf. Sci. | 5 |
| 2025 | Speech2Face3D: A Two-Stage Transfer-Learning Framework for Speech-Driven 3D Facial AnimationabstractABSTRACT High‐fidelity, speech‐driven 3D facial animation is crucial for immersive applications and virtual avatars. Nevertheless, advancement is impeded by two principal challenges: (1) a lack of high‐quality 3D data, and (2) inadequate modelling of the multi‐scale characteristics of speech signals. In this paper, we present Speech2Face3D, a novel two‐stage transfer‐learning framework that pretrains on large‐scale pseudo‐3D facial data derived from 2D videos and subsequently finetunes on smaller yet high‐fidelity 3D datasets. This design leverages the richness of easily accessible 2D resources while mitigating reconstruction noise through a simple temporal smoothing step. Our approach further introduces a Multi‐Scale Hierarchical Audio Encoder to capture subtle phoneme transitions, mid‐range prosody, and longer‐range emotional cues. Extensive experiments on public 3D benchmarks demonstrate that our method achieves state‐of‐the‐art performance on lip synchronization, expression fidelity, and temporal coherence metrics. Qualitative user evaluations validate these quantitative improvements. Speech2Face3D is a robust and scalable framework for utilizing extensive 2D data to generate precise and realistic 3D facial animations only based on speech. Liming Pang, Guixuan Zhang, Shuwu Zhang |
IET Image Process. | 4 |
| 2025 | Weakening the Dominant Role of Text: CMOSI Dataset and Multimodal Semantic Enhancement NetworkabstractMultimodal sentiment analysis (MSA) is important for quickly and accurately understanding people's attitudes and opinions about an event. However, existing sentiment analysis methods suffer from the dominant contribution of text modality in the dataset; this is called text dominance. In this context, we emphasize that weakening the dominant role of text modality is important for MSA tasks. To solve the above two problems, from the perspective of datasets, we first propose the Chinese multimodal opinion-level sentiment intensity (CMOSI) dataset. Three different versions of the dataset were constructed: manually proofreading subtitles, generating subtitles using machine speech transcription, and generating subtitles using human cross-language translation. The latter two versions radically weaken the dominant role of the textual model. We randomly collected 144 real videos from the Bilibili video site and manually edited 2557 clips containing emotions from them. From the perspective of network modeling, we propose a multimodal semantic enhancement network (MSEN) based on a multiheaded attention mechanism by taking advantage of the multiple versions of the CMOSI dataset. Experiments with our proposed CMOSI show that the network performs best with the text-unweakened version of the dataset. The loss of performance is minimal on both versions of the text-weakened dataset, indicating that our network can fully exploit the latent semantics in nontext patterns. In addition, we conducted model generalization experiments with MSEN on MOSI, MOSEI, and CH-SIMS datasets, and the results show that our approach is also very competitive and has good cross-language robustness. Ming Yan 0005, Guangzhe Zhao, Guixuan Zhang, Shuwu Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion
Zeyu Zhao 0005, Nan Gao 0001, Guixuan Zhang, Jie Liu 0028, Shuwu Zhang |
MMAsia | 4 |
| 2024 | Degradation regression with uncertainty for blind super-resolution
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang |
Neurocomputing | 2 |
| 2023 | Robust Texture-Aware Local Adaptive Image Watermarking With Perceptual GuaranteeabstractWatermarking involves embedding a watermark in an image and later extracting it to prove the image’s copyright. In most cases, a complete image contains both smooth and textured regions. As a rule of thumb, the visual quality of an image with a watermark embedded in its textured regions is better than that of the same image with a watermark in smooth regions. This paper, by taking advantage of the fact, proposes a texture-aware local adaptive watermarking algorithm to maximize the watermark’s robustness while maintaining its imperceptibility. To identify textured regions in an image, we introduce the texture value, an efficient and proper metric of the richness of image texture. It combines the texture correlation of the AC coefficients, the luminance masking of the DC coefficient, and the distribution of image texture. A watermark is embedded adaptively into multiple non-overlapping textured regions of an image under the specified SSIM condition. Its adaptiveness comes from a novel texture-aware adaptive parameter model derived by multivariate regression analysis. Correct extraction of watermarks from multiple textured regions can be done by the cooperation of embedding and extraction strategies, with the assistance of RS-based watermark coding model. They allow for greater robustness, faster extraction, and adjustable watermark capacity. The simulation experiments on 100 images demonstrate that our proposed algorithm outperforms state-of-the-art algorithms with respect to imperceptibility, robustness, and adaptability. Hu Guan, Jie Liu 0028, Shuwu Zhang, Baoning Niu, Guixuan Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Heterogeneous Avatar Synthesis Based on Disentanglement of Topology and Rendering
Guixuan Zhang, Shuwu Zhang |
ACCV (4) | 3 |
| 2022 | From general to specific: Online updating for blind super-resolution
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang |
Pattern Recognit. | 2 |
| 2021 | Approaching the Limit of Image Rescaling via Flow Guidance
Guixuan Zhang, Zhengxiong Luo 0001, Jie Liu 0028, Shuwu Zhang |
BMVC | 2 |
| 2021 | Learning to predict more accurate text instances for scene text detection
Jie Liu 0028, Guixuan Zhang, Yang Zheng 0002, Shuwu Zhang |
Neurocomputing | 3 |
| 2020 | IBN-STR: A Robust Text Recognizer for Irregular Text in Natural ScenesabstractAlthough text recognition methods based on deep neural networks have promising performance, there are still challenges due to the variety of text styles, perspective distortion, text with large curvature, and so on. To obtain a robust text recognizer, we have improved the performance from two aspects: data aspect and feature representation aspect. In terms of data, we transform the input images into S-shape distorted images in order to increase the diversity of training data. Besides, we explore the effects of different training data. In terms of feature representation, the combination of instance normalization and batch normalization improves the model's capacity and generalization ability. This paper proposes a robust scene text recognizer IBN-STR, which is an attention-based model. Through extensive experiments, the model analysis and comparison have been carried out from the aspects of data and feature representation, and the effectiveness of IBN-STR on both regular and irregular text instances has been verified. Furthermore, IBN-STR is an end-to-end recognition system that can achieve state-of-the-art performance. Jie Liu 0028, Guixuan Zhang, Shuwu Zhang |
ICPR | 3 |
| 2020 | Single shot multi-oriented text detection based on local and non-local features
Jie Liu 0028, Shuwu Zhang, Guixuan Zhang, Yang Zheng 0002 |
Int. J. Document Anal. Recognit. | 4 |
| 2018 | Aspect-Level Sentiment Classification with Conv-Attention Mechanism
Jie Liu 0028, Guixuan Zhang, Shuwu Zhang |
ICONIP (4) | 3 |
| 2017 | Region based image retrieval with query-adaptive feature fusionabstractRecently, image representation based on convolutional neural network (CNN) becomes more popular than SIFT based feature, such as Fisher vector (FV). However, which of the two works better for image retrieval is not entirely clear yet. In this paper, we propose to fuse CNN and FV to incorporate the advantages of both features for image retrieval. We extract CNN feature and FV from multi-scale regions, which makes the representation more robust to image noise. Then a query-adaptive feature fusion method is proposed, which is used jointly with 2-D inverted index under the framework of bag-of-words. Moreover, we make an evaluation of different CNN feature extraction methods for the region based method. Extensive experiments on four benchmark datasets demonstrate the effectiveness of our method with efficiency in both time cost and memory usage. Guixuan Zhang, Shuwu Zhang, Hu Guan, Fangxin Wang 0002 |
ICIP | 1 |
| 2017 | SIFT Matching with CNN Evidences for Particular Object Retrieval
Guixuan Zhang, Shuwu Zhang, Wanchun Wu |
Neurocomputing | 1 |
| 2016 | Region matching and similarity enhancing for image retrievalabstractMany image retrieval systems adopt the bag-of-words model and rely on matching of local descriptors. However, these descriptors of keypoints, such as SIFT, may lead to false matches, since they do not consider the contextual information of the keypoints. In this paper, we incorporate the cues of meaningful regions where local descriptors are extracted. We describe a matching region estimation (MRE) method to find appropriate matching regions for local descriptor matching pairs. Then the region matching quality is evaluated and the true matched regions will enhance the similarity of local descriptors. Consequently, the image retrieval accuracy can be improved. Extensive experiments on benchmark datasets show the effectiveness of our method and our result compares favorably with the state-of-the-art. Guixuan Zhang, Shuwu Zhang, Hu Guan, Qin-Zhen Guo |
ICASSP | 1 |
| 2016 | Adaptive bit allocation product quantization
Qin-Zhen Guo, Shuwu Zhang, Guixuan Zhang |
Neurocomputing | 4 |
| 2015 | Transmitting informative components of fisher codes for mobile visual searchabstractExisting techniques usually adopt compact descriptors such as Fisher vector for mobile visual search, since compact descriptors are memory-efficient and suitable for fast transmission. In common Fisher vector methods, in order to make the size of image representations small enough for efficient transmission, only a small number of visual words are used. However, this choice usually sacrifices the search accuracy. In this paper, a Soft-Assignment Adjusting approach is proposed to just select informative components of descriptors for query. With this method, we can adopt more visual words to improve accuracy, while the memory usage is still low. Furthermore, efficient bitrate scalable codes are proposed in order to accommodate the network bandwidth variation. Experiments performed on benchmark datasets show that our proposed approach outperforms the state-of-the-art methods for mobile visual search. Guixuan Zhang, Shuwu Zhang, Qin-Zhen Guo |
ICASSP | 1 |