VLDB 2026 Research / reviewers in the wild / expert
Ziqi Gu
dblp:256/5617
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0009-0003-5938-2621ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Noisy Candidates to Reliable Grounding in Weakly-Supervised Referring Expression ComprehensionabstractWeakly-Supervised Referring Expression Comprehension (WREC) aims to ground natural language expressions in image regions, using only image-text pairs without bounding-box annotations. However, the absence of explicit localization supervision results in noisy training signals and unreliable candidate selection. We observe that modern query-based detectors exhibit a high-recall, low-precision behavior in WREC: although the top-ranked query often localizes inaccurately, the ground-truth region is frequently recalled among high-confidence candidates. Motivated by this observation, we propose a Noisy-to-Reliable Grounding (NRG) framework that progressively transforms noisy high-recall candidates into reliable grounding supervision. Given an paired image-text data, we leverage a large Vision-Language Model (VLM) to generate confidence-aware soft pseudo labels, which provide robust semantic and spatial guidance under weak supervision. To disentangle true positives from noisy candidates, we introduce a contrastive candidate mining module that jointly exploits the detector’s prediction and VLM-derived cues to identify positive and hard negative queries, progressively enhancing grounding discriminability. Furthermore, the confidence-aware pseudo labels are integrated to construct a regression objective, enabling effective localization learning through low-rank adaptation of a pretrained open-vocabulary detector. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate that the proposed NRG consistently outperforms existing WREC methods. Ziqi Gu, Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu |
ICMR | 2 |
| 2026 | Boosting Few-Shot Continual Learning via Self-Adaptive EvolutionabstractFew-shot continual learning (FSCL) has attracted increasing attention for real-world applications, where models must continuously adapt to new classes with only a few labeled samples while retaining prior knowledge. These abilities are essential in dynamic environments where data availability is often sparse and nonstationary. However, traditional FSCL methods are largely confined to closed data spaces, which limits their generalizability when diverse and evolving distributions are involved. Inspired by the paradigm of human lifelong learning, we propose a new self-adaptive evolution framework for FSCL that enables continuous interaction with and adaptation to external environments. To exploit latent knowledge in large-scale models, we use an adaptive diffusion-based generator that not only implicitly captures the distribution of new few-shot samples but also produces more high-quality samples. To mitigate the inevitable variability in generation quality, we also use a reinforced sample selection module, comprising a generated sample explorer and a selection evaluator, which explicitly guides the retained distributions toward alignment with the large-scale models. Integrated with the continual model, these components are optimized in an iterative self-adaptive evolution framework, ensuring stable knowledge retention while improving adaptability to newly emerging classes. We validate our approach through experiments on three benchmarks, revealing its effectiveness in exploiting external distributions and achieving notable performance improvements. Ziqi Gu, Chunyan Xu, Yuanzhi Wang, Cao Han, Di Xia, Zhen Cui 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Multi-clue Consistency Learning to Bridge Gaps Between General and Oriented Object in Semi-supervised DetectionabstractWhile existing semi-supervised object detection (SSOD) methods perform well in general scenes, they encounter challenges in handling oriented objects in aerial images. We experimentally find three gaps between general and oriented object detection in semi-supervised learning: 1) Sampling inconsistency: the common center sampling is not suitable for oriented objects with larger aspect ratios when selecting positive labels from labeled data. 2) Assignment inconsistency: balancing the precision and localization quality of oriented pseudo-boxes poses greater challenges which introduces more noise when selecting positive labels from unlabeled data. 3) Confidence inconsistency: there exists more mismatch between the predicted classification and localization qualities when considering oriented objects, affecting the selection of pseudo-labels. Therefore, we propose a Multi-clue Consistency Learning (MCL) framework to bridge gaps between general and oriented objects in semi-supervised detection. Specifically, considering various shapes of rotated objects, the Gaussian Center Assignment is specially designed to select the pixel-level positive labels from labeled data. We then introduce the Scale-aware Label Assignment to select pixel-level pseudo-labels instead of unreliable pseudo-boxes, which is a divide-and-rule strategy suited for objects with various scales. The Consistent Confidence Soft Label is adopted to further boost the detector by maintaining the alignment of the predicted results. Comprehensive experiments on DOTA-v1.5 and DOTA-v1.0 benchmarks demonstrate that our proposed MCL can achieve state-of-the-art performance in the semi-supervised oriented object detection task. Chunyan Xu, Xiang Li 0041, YuXuan Li, Ziqi Gu, Zhen Cui 0001 |
AAAI | 6 |
| 2025 | CLIP-driven Few-Shot Continual LearningabstractFew-Shot continual learning (FSCL) has garnered significant attention and has made notable progress in recent years. Drawing inspiration from the continuous interaction of humans with their environment in the lifelonglearning [1], we propose a novel CLIP-driven Few-Shot Continual Learning (C-FSCL) framework that aims to progressively enhance the continual network by leveraging the capabilities of a foundational CLIP model. When the continual model encounters new-coming data, we introduce a hierarchical feature-aware alignment module to better generalize knowledge gained from the foundational CLIP model, allowing it to apply learned concepts to new tasks more effectively. Recognizing the instability of inter-class structural relationships, which leads to catastrophic forgetting, we further introduce a relation-aware structure alignment module to align the consistent relationships from the CLIP model, thereby enhancing the reliability of the continual model’s knowledge structure. Experiments demonstrate the feasibility of CLIP-driven FSCL task and its remarkable performance on three benchmarks. Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
ICME | 1 |
| 2025 | Learn and Ensemble Bridge Adapters for Multi-domain Task Incremental LearningabstractMulti-domain task incremental learning (MTIL) demands models to master domain-specific expertise while preserving generalization capabilities.
Inspired by human lifelong learning, which relies on revisiting, aligning, and integrating past experiences, we propose a Learning and Ensembling Bridge Adapters (LEBA) framework.
To facilitate cohesive knowledge transfer across domains, specifically, we propose a continuous-domain bridge adaptation module, leveraging the distribution transfer capabilities of Schrödinger bridge for stable progressive learning.
To strengthen memory consolidation, we further propose a progressive knowledge ensemble strategy that revisits past task representations via a diffusion model and dynamically integrates historical adapters.
For efficiency, LEBA maintains a compact adapter pool through similarity-based selection and employs learnable weights to align replayed samples with current task semantics.
Together, these components effectively mitigate catastrophic forgetting and enhance generalization across tasks.
Extensive experiments across multiple benchmarks validate the effectiveness and superiority of LEBA over state-of-the-art methods. Ziqi Gu, Chunyan Xu, Xin Liu 0011, Yide Qiu, Zhen Cui 0001 |
NeurIPS | 1 |
| 2025 | UniHG: A Large-scale Universal Heterogeneous Graph Dataset and Benchmark for Representation Learning and Cross-Domain TransferringabstractIrregular data in the real world are usually organized as heterogeneous graphs consisting of multiple types of nodes and edges. However, current heterogeneous graph research confronts three fundamental challenges: i) Benchmark Deficiency, ii) Semantic Disalignment, and iii) Propagation Degradation. In this paper, we construct a large-scale, universal, and joint multi-domain heterogeneous graph dataset named UniHG to facilitate heterogeneous graph representation learning and cross-domain knowledge mining. Overall, UniHG contains 77.31 million nodes and 564 million directed edges with thousands of labels and attributes, which is currently the largest universal heterogeneous graph dataset available to the best of our knowledge. To perform effective learning and provide comprehensively benchmarks on UniHG , two key measures are taken, including i) the semantic alignment strategy for multi-attribute entities, which projects the feature description of multi-attribute nodes and edges into a common embedding space to facilitate information aggregation; ii) proposing the novel Heterogeneous Graph Decoupling (HGD) framework with a specifically designed Anisotropy Feature Propagation (AFP) module for learning effective multi-hop anisotropic propagation kernels. These two strategies enable efficient information propagation among a tremendous number of multi-attribute entities and meanwhile mine multi-attribute association adaptively through the multi-hop aggregation in large-scale heterogeneous graphs. Comprehensive benchmark results demonstrate that our model significantly outperforms existing methods with an accuracy improvement of 28.93\%. And the UniHG can facilitate downstream tasks, achieving an NDCG@20 improvement rate of 11.48\% and 11.71\%. The UniHG dataset and benchmark codes have been released at https://anonymous.4open.science/r/UniHG-AA78. Yide Qiu, Tong Zhang 0021, Shaoxiang Ling, Xing Cai, Ziqi Gu, Zhen Cui 0001 |
NeurIPS | 5 |
| 2025 | Adaptive exploration for few-shot incremental learning
Cao Han, Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Semi-Supervised Building Footprint Extraction Using Debiased Pseudo-LabelsabstractAccurate extraction of building footprints from satellite imagery is of high value. Currently, deep learning methods are predominant in this field due to their powerful representation capabilities. However, they generally require extensive pixel-wise annotations, which constrains their practical application. Semi-supervised learning (SSL) significantly mitigates this requirement by leveraging large volumes of unlabeled data for model self-training (ST), thus enhancing the viability of building footprint extraction. Despite its advantages, SSL faces a critical challenge: the imbalanced distribution between the majority background class and the minority building class, which often results in model bias toward the background during training. To address this issue, this article introduces a novel method called DeBiased matching (DBMatch) for semi-supervised building footprint extraction. DBMatch comprises three main components: 1) a basic supervised learning module (SUP) that uses labeled data for initial model training; 2) a classical weak-to-strong ST module that generates pseudo-labels from unlabeled data for further model ST; and 3) a novel logit debiasing (LDB) module that calculates a global logit bias between building and background, allowing for dynamic pseudo-label calibration. To verify the effectiveness of the proposed DBMatch, extensive experiments are performed on three public building footprint extraction datasets covering six global cities in SSL setting. The experimental results demonstrate that our method significantly outperforms some advanced SSL methods in semi-supervised building footprint extraction. Our codes will be publicly provided athttps://github.com/zhu-xlab/SSL_Buildings. Wei Huang 0068, Ziqi Gu, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Learning Building Energy Efficiency with Semantic AttributesabstractNon-intrusive estimation of building energy efficiency has profound applications in advancing sustainability in the built environment. Recent studies often focus on predicting energy performance alone, neglecting the interplay between the performance and related building semantics. This paper investigates whether incorporating semantic attributes benefits energy efficiency estimation. We develop a neural network to estimate energy efficiency, with building age and usage type as additional supervision for multi-task learning. The neural network processes both aerial imagery and airborne LiDAR data to classify buildings as energy-efficient or inefficient. Our results demonstrate the effectiveness of the superimposed semantics, particularly with building age. With the multi-task model achieving a 63.78% F1 score and outperforming that supervised solely with energy efficiency by 2.86%, this paper reveals the potential of integrating semantic attributes in modeling building energy performance. Zhaiyu Chen, Ziqi Gu, Yilei Shi, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2024 | Building Attributes Recognition with Noisy and Incomplete LabelsabstractRecognizing building attributes from remote sensing images is crucial for various applications. Recent developments in deep learning have demonstrated promising results in identifying these attributes. Nonetheless, a major challenge is the requirement for extensive and accurate building attribute data. Two primary data sources are commonly considered: Open-StreetMap (OSM), which offers global building information but often lacks completeness and correctness, and cadastral data, known for its high quality but typically restricted to certain areas. These two sources enable comparison between deep learning models trained on noisy and incomplete OSM data and those trained on accurate and complete cadastral data. In this work, comprehensive experiments on buildings in Bavaria, Germany, are conducted, covering diverse attributes such as footprints, use, and height. A large building dataset with corresponding building attribute labels from OSM and cadastral data is created, with OSM data featuring varying levels of incompleteness and noise for different attributes and cadastral data serving as ground truth. Moreover, we evaluate the effectiveness of several prevailing methods designed to handle noisy and incomplete labels, assessing their applicability to real-world scenarios with incomplete and noisy OSM labels. Ziqi Gu, Zhaiyu Chen, Yilei Shi, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2024 | MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image GenerationabstractRecently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general images in terms of scale and perspective remains a formidable challenge due to the lack of a comprehensive remote sensing image generation dataset with various modalities, ground sample distances (GSD), and scenes. In this paper, we propose a Multi-modal, Multi-GSD, Multi-scene Remote Sensing (MMM-RS) dataset and benchmark for text-to-image generation in diverse remote sensing scenarios. Specifically, we first collect nine publicly available RS datasets and conduct standardization for all samples. To bridge RS images to textual semantic information, we utilize a large-scale pretrained vision-language model to automatically output text prompts and perform hand-crafted rectification, resulting in information-rich text-image pairs (including multi-modal images). In particular, we design some methods to obtain the images with different GSD and various environments (e.g., low-light, foggy) in a single sample. With extensive manual screening and refining annotations, we ultimately obtain a MMM-RS dataset that comprises approximately 2.1 million text-image pairs. Extensive experimental results verify that our proposed MMM-RS dataset allows off-the-shelf diffusion models to generate diverse RS images across various modalities, scenes, weather conditions, and GSD. The dataset is available at https://github.com/ljl5261/MMM-RS. Jialin Luo, Yuanzhi Wang, Ziqi Gu, Yide Qiu, Shuaizhen Yao, Fuyun Wang, Chunyan Xu, Zhen Cui 0001 |
NeurIPS | 3 |
| 2023 | Few-shot Continual Infomax LearningabstractFew-shot continual learning is the ability to continually train a neural network from a sequential stream of few-shot data. In this paper, we propose a Few-shot Continual Infomax Learning (FCIL) framework that makes a deep model to continually/incrementally learn new concepts from few labeled samples, relieving the catastrophic forgetting of past knowledge. Specifically, inspired by the theoretical definition of transfer entropy, we introduce a feature embedding infomax to effectively perform the few-shot learning, which can transfer the strong encoding capability of the base network to learn the feature embedding of these novel classes by maximizing the mutual information of different-level feature distributions. Further, considering that the learned knowledge in the human brain is a generalization of actual information and exists in a certain relational structure, we perform continual structure infomax learning to relieve the catastrophic forgetting problem in the continual learning process. The information structure of this learned knowledge can be preserved through maximizing the mutual information across these continual-changing relations of inter-classes. Comprehensive evaluations on CIFAR100, miniImageNet, and CUB200 datasets demonstrate the superiority of our FCIL when compared against state-of-the-art methods on the few-shot continual learning task. Ziqi Gu, Chunyan Xu, Jian Yang 0003, Zhen Cui 0001 |
ICCV | 1 |
| 2023 | Grassmann Graph Embedding for Few-Shot Class Incremental Learning
Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
PRCV (8) | 1 |