VLDB 2026 Research / reviewers in the wild / expert
Zhiyong Gan
dblp:267/9223
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An alignment-error-free framework for end-to-end table recognition
Fan Yang 0082, Ling Deng, Zhiyong Gan, Shuangping Huang, Tianshui Chen |
Expert Syst. Appl. | 3 |
| 2026 | Mixture of knowledge from multiple pre-trained foundation models for few-shot recognition
Junxi Chen, Zhiyong Gan, Guangxing Wu |
Neurocomputing | 2 |
| 2026 | A text-only weakly supervised learning framework for text spotting via text-to-polygon generator
Zhiyong Gan, Ling Deng, Shuaicheng Niu, Zhenghua Peng, Shuangping Huang |
Pattern Recognit. | 2 |
| 2025 | ReplayCAD: Generative Diffusion Replay for Continual Anomaly DetectionabstractContinual Anomaly Detection (CAD) enables anomaly detection models in learning new classes while preserving knowledge of historical classes. CAD faces two key challenges: catastrophic forgetting and segmentation of small anomalous regions. Existing CAD methods store image distributions or patch features to mitigate catastrophic forgetting, but they fail to preserve pixel-level detailed features for accurate segmentation. To overcome this limitation, we propose ReplayCAD, a novel diffusion-driven generative replay framework that replay high-quality historical data, thus effectively preserving pixel-level detailed features. Specifically, we compress historical data by searching for a class semantic embedding in the conditional space of the pre-trained diffusion model, which can guide the model to replay data with fine-grained pixel details, thus improving the segmentation performance. However, relying solely on semantic features results in limited spatial diversity. Hence, we further use spatial features to guide data compression, achieving precise control of sample space, thereby generating more diverse data. Our method achieves state-of-the-art performance in both classification and segmentation, with notable improvements in segmentation: 11.5% on VisA and 8.1% on MVTec. Our source code is available at https://github.com/HULEI7/ReplayCAD. Lei Hu 0012, Zhiyong Gan, Ling Deng, Jinglin Liang 0001, Lingyu Liang, Shuangping Huang, Tianshui Chen |
IJCAI | 2 |
| 2025 | Optimal Feature Embedding for Document Large Visual Language ModelabstractDocument Large Vision Language Models excel in document-centric tasks and have become a key focus of research. Existing frameworks embed features from a lightweight, document-specific encoder into the first layer of a general-purpose Vision Language Model (VLM). However, this introduces a feature mismatch problem. VLMs typically consist of many stacked layers, with the feature hierarchy becoming increasingly abstract at higher layers. Specifically, the first-layer feature in a VLM is token-level, whereas the feature from the encoder is task-level, resulting in a mismatch. Consequently, it is crucial to identify an optimal layer within the VLM for embedding the encoder's features. Inspired by physics, we reformulate the search for the optimal embedding as a problem of finding the shortest time curve. Leveraging the properties of the shortest time curve, we theoretically derive a task-agnostic proxy score that requires only partial training and propose our searching framework, Brac4VLM. Our theoretical derivation shows that Brac4VLM reduces search time by 97.8% compared to brute-force methods. Experimental results further demonstrate that Brac4VLM identifies embedding points that closely align with the true optima. Moreover, the DocVLM with the optimal embedding position identified achieves state-of-the-art performance across various document-centric tasks. Codes: https://github.com/MaxKinny/Brac4VLM. Fan Yang 0082, Ling Deng, Zhiyong Gan, Qisheng He, Yuanbo Fang, Xiangmin Xu 0001, Shuangping Huang, Tianshui Chen |
ACM Multimedia | 3 |
| 2024 | Out-of-Distribution Detection by Principal Component CorrespondenceabstractOut-of-distribution (OOD) detection is vital for the safe application of intelligent systems in real-world scenarios. This paper proposes an enhancement to OOD detection by leveraging the consistency in cognition between two models, both pretrained on in-distribution (ID) data. Specifically, for a given test sample, we first apply Principal Component Analysis (PCA)-based projection on the feature vectors from each model. These obtained feature vectors (with correlation between dimensions decoupled by PCA projection) are then aligned using a multiple linear mapping, which is fitted using the least squares method on the training data. We hypothesize that the regression error for OOD data will be larger than that for ID data, making it a useful metric for OOD detection. Our experimental results demonstrate the effectiveness of this method. When combined with existing robust baselines, our approach achieves state-of-the-art performance in OOD detection. Xiaoyuan Guan, Zhiyong Gan, Ling Deng, Jiankang Chen, Shenshen Bu, Chunliang Zhao, Jianfang Hu, Wei-Shi Zheng 0001 |
ICME | 2 |
| 2024 | FodFoM: Fake Outlier Data by Foundation Models Creates Stronger Visual Out-of-Distribution DetectorabstractOut-of-Distribution (OOD) detection is crucial when deploying machine learning models in open-world applications. The core challenge in OOD detection is mitigating the model's overconfidence on OOD data. While recent methods using auxiliary outlier datasets or synthesizing outlier features have shown promising OOD detection performance, they are limited due to costly data collection or simplified assumptions. In this paper, we propose a novel OOD detection framework FodFoM that innovatively combines multiple foundation models to generate two types of challenging fake outlier images for classifier training. The first type is based on BLIP-2's image captioning capability, CLIP's vision-language knowledge, and Stable Diffusion's image generation ability. Jointly utilizing these foundation models constructs fake outlier images which are semantically similar to but different from in-distribution (ID) images. For the second type, GroundingDINO's object detection ability is utilized to help construct pure background images by blurring foreground ID objects in ID images. The proposed framework can be flexibly combined with multiple existing OOD detection methods. Extensive empirical evaluations show that image classifiers with the help of constructed fake images can more accurately differentiate real OOD image from ID ones. New state-of-the-art OOD detection performance is achieved on multiple benchmarks. The code is available at https://github.com/Cverchen/ACMMM2024-FodFoM. Jiankang Chen, Ling Deng, Zhiyong Gan, Wei-Shi Zheng 0001 |
ACM Multimedia | 3 |
| 2024 | On defect restricted matching extension graphs
Zhiyong Gan |
Discret. Appl. Math. | 1 |
| 2021 | Hamiltonian and long cycles in bipartite graphs with connectivity
Zhiyong Gan |
Discret. Appl. Math. | 1 |