VLDB 2026 Research / reviewers in the wild / expert
Guangtao Zheng
dblp:178/7288
· DBLP profile ↗
16ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-1287-4931ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Computer networks · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal BiasabstractLarge vision-language models such as CLIP have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models often develop multimodal spurious biases, the undesirable tendency to rely on spurious features. For example, CLIP may infer object types in images based on frequently co-occurring backgrounds rather than the object's core features. This bias significantly impairs the robustness of pre-trained CLIP models on out-of-distribution data, where such cross-modal associations no longer hold. Existing methods for mitigating multimodal spurious bias typically require fine-tuning on downstream data or prior knowledge of the bias, which undermines the out-of-the-box usability of CLIP. In this paper, we first theoretically analyze the impact of multimodal spurious bias in zero-shot classification. Based on this insight, we propose Spuriousness-Aware Guided Exploration (SAGE), a simple and effective method that mitigates spurious bias via guided prompt selection. SAGE requires no training, fine-tuning, or external annotations. It explores on a space of prompt templates and selects the prompts that induces the largest semantic separation between classes, thereby improving worst-group robustness. Extensive experiments on four real-world benchmark datasets and five popular backbone models demonstrate that SAGE consistently improves zero-shot performance and generalization, outperforming previous zero-shot approaches without any external knowledge or model updates. Wenqian Ye, Di Wang 0053, Guangtao Zheng, Bohan Liu 0008, Aidong Zhang 0001 |
AAAI | 3 |
| 2026 | MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMsabstractSpurious bias, a tendency to exploit spurious correlations between superficial input attributes and prediction targets, has revealed a severe robustness pitfall in classical machine learning problems. Multimodal Large Language Models (MLLMs), which leverage pretrained vision and language models, have recently demonstrated strong capability in joint vision-language understanding. However, both the presence and severity of spurious biases in MLLMs remain poorly understood. In this work, we address this gap by analyzing the spurious biases in the multimodal setting and uncovering the specific inference-time data patterns that can manifest this problem. To support this analysis, we introduce MM-SpuBench, a comprehensive, human-verified benchmark dataset consisting of image-class pairs annotated with core and spurious attributes, grounded in our taxonomy of nine distinct types of spurious correlations. The benchmark is constructed using human-interpretable attribute information to capture a wide range of spurious patterns reflective of real-world knowledge. Leveraging this benchmark, we conduct a comprehensive evaluation of the state-of-the-art open-source and proprietary MLLMs with both standard accuracy and the proposed Conditional Generation Likelihood Advantage (CGLA). Our findings highlight the persistence of reliance on spurious correlations and the difficulty of mitigation on our benchmark. We hope this work can inspire new technical strides to mitigate these biases. Our benchmark is publicly available at https://huggingface.co/datasets/mmbench/MM-SpuBench. Wenqian Ye, Bohan Liu 0008, Guangtao Zheng, Di Wang 0053, Yunsheng Ma, Bolin Lai, James M. Rehg, Aidong Zhang 0001 |
KDD (1) | 3 |
| 2025 | NeuronTune: Towards Self-Guided Spurious Bias MitigationabstractDeep neural networks often develop spurious bias, reliance on correlations between non-essential features and classes for predictions. For example, a model may identify objects based on frequently co-occurring backgrounds rather than intrinsic features, resulting in degraded performance on data lacking these correlations. Existing mitigation approaches typically depend on external annotations of spurious correlations, which may be difficult to obtain and are not relevant to the spurious bias in a model. In this paper, we take a step towards self-guided mitigation of spurious bias by proposing NeuronTune, a post hoc method that directly intervenes in a model’s internal decision process. Our method probes in a model’s latent embedding space to identify and regulate neurons that lead to spurious prediction behaviors. We theoretically justify our approach and show that it brings the model closer to an unbiased one. Unlike previous methods, NeuronTune operates without requiring spurious correlation annotations, making it a practical and effective tool for improving model robustness. Experiments across different architectures and data modalities demonstrate that our method significantly mitigates spurious bias in a self-guided way. Guangtao Zheng, Wenqian Ye, Aidong Zhang 0001 |
ICML | 1 |
| 2025 | ShortcutProbe: Probing Prediction Shortcuts for Learning Robust ModelsabstractDeep learning models often achieve high performance by inadvertently learning spurious correlations between targets and non-essential features. For example, an image classifier may identify an object via its background that spuriously correlates with it. This prediction behavior, known as spurious bias, severely degrades model performance on data that lacks the learned spurious correlations. Existing methods on spurious bias mitigation typically require a variety of data groups with spurious correlation annotations called group labels. However, group labels require costly human annotations and often fail to capture subtle spurious biases such as relying on specific pixels for predictions. In this paper, we propose a novel post hoc spurious bias mitigation framework without requiring group labels. Our framework, termed ShortcutProbe, identifies prediction shortcuts that reflect potential non-robustness in predictions in a given model's latent space. The model is then retrained to be invariant to the identified prediction shortcuts for improved robustness. We theoretically analyze the effectiveness of the framework and empirically demonstrate that it is an efficient and practical tool for improving a model's robustness to spurious bias on diverse datasets. Guangtao Zheng, Wenqian Ye, Aidong Zhang 0001 |
IJCAI | 1 |
| 2025 | Improving Group Robustness on Spurious Correlation via Evidential AlignmentabstractDeep neural networks often learn and rely on spurious correlations, i.e., superficial associations between non-causal features and the targets. For instance, an image classifier may identify camels based on the desert backgrounds. While it can yield high overall accuracy during training, it degrades generalization on more diverse scenarios where such correlations do not hold. This problem poses significant challenges for out-of-distribution robustness and trustworthiness. Existing methods typically mitigate this issue by using external group annotations or auxiliary deterministic models to learn unbiased representations. However, such information is costly to obtain, and deterministic models may fail to capture the full spectrum of biases learned by the models. To address these limitations, we propose Evidential Alignment, a novel framework that leverages uncertainty quantification to understand the behavior of the biased models without requiring group annotations. By quantifying the evidence of model prediction with second-order risk minimization and calibrating the biased models with the proposed evidential calibration technique, Evidential Alignment identifies and suppresses spurious correlations while preserving core features. We theoretically justify the effectiveness of our method as capable of learning the patterns of biased models and debiasing the model without requiring any spurious correlation annotations. Empirical results demonstrate that our method significantly improves group robustness across diverse architectures and data modalities, providing a scalable and principled solution to spurious correlations. Wenqian Ye, Guangtao Zheng, Aidong Zhang 0001 |
KDD (2) | 2 |
| 2025 | Rectifying Shortcut Behaviors in Preference-based Reward LearningabstractIn reinforcement learning from human feedback, preference-based reward models play a central role in aligning large language models to human-aligned behavior. However, recent studies show that these models are prone to reward hacking and often fail to generalize well due to over-optimization. They achieve high reward scores by exploiting shortcuts, that is, exploiting spurious features (e.g., response verbosity, agreeable tone, or sycophancy) that correlate with human preference labels in the training data rather than genuinely reflecting the intended objectives. In this paper, instead of probing these issues one at a time, we take a broader view of the reward hacking problem as shortcut behaviors and introduce a principled yet flexible approach to mitigate shortcut behaviors in preference-based reward learning. Inspired by the invariant theory in the kernel perspective, we propose Preference-based Reward Invariance for Shortcut Mitigation (PRISM), which learns group-invariant kernels with feature maps in a closed-form learning objective. Experimental results in several benchmarks show that our method consistently improves the accuracy of the reward model on diverse out-of-distribution tasks and reduces the dependency on shortcuts in downstream policy models, establishing a robust framework for preference-based alignment. Wenqian Ye, Guangtao Zheng, Aidong Zhang 0001 |
NeurIPS | 2 |
| 2024 | AdvST: Revisiting Data Augmentations for Single Domain GeneralizationabstractSingle domain generalization (SDG) aims to train a robust model against unknown target domain shifts using data from a single source domain. Data augmentation has been proven an effective approach to SDG. However, the utility of standard augmentations, such as translate, or invert, has not been fully exploited in SDG; practically, these augmentations are used as a part of a data preprocessing procedure. Although it is intuitive to use many such augmentations to boost the robustness of a model to out-of-distribution domain shifts, we lack a principled approach to harvest the benefit brought from multiple these augmentations. Here, we conceptualize standard data augmentations with learnable parameters as semantics transformations that can manipulate certain semantics of a sample, such as the geometry or color of an image. Then, we propose Adversarial learning with Semantics Transformations (AdvST) that augments the source domain data with semantics transformations and learns a robust model with the augmented data. We theoretically show that AdvST essentially optimizes a distributionally robust optimization objective defined on a set of semantics distributions induced by the parameters of semantics transformations. We demonstrate that AdvST can produce samples that expand the coverage on target domain data. Compared with the state-of-the-art methods, AdvST, despite being a simple method, is surprisingly competitive and achieves the best average SDG performance on the Digits, PACS, and DomainNet datasets. Our code is available at https://github.com/gtzheng/AdvST. Guangtao Zheng, Mengdi Huai, Aidong Zhang 0001 |
AAAI | 1 |
| 2024 | Benchmarking Spurious Bias in Few-Shot Image Classifiers
Guangtao Zheng, Wenqian Ye, Aidong Zhang 0001 |
ECCV (80) | 1 |
| 2024 | Learning Robust Classifiers with Self-Guided Spurious Correlation Mitigation
Guangtao Zheng, Wenqian Ye, Aidong Zhang 0001 |
IJCAI | 1 |
| 2024 | Spuriousness-Aware Meta-Learning for Learning Robust ClassifiersabstractSpurious correlations are brittle associations between certain attributes of inputs and target variables, such as the correlation between an image background and an object class. Deep image classifiers often leverage them for predictions, leading to poor generalization on the data where the correlations do not hold. Mitigating the impact of spurious correlations is crucial towards robust model generalization, but it often requires annotations of the spurious correlations in data -- a strong assumption in practice. In this paper, we propose a novel learning framework based on meta-learning, termed SPUME -- SPUriousness-aware MEta-learning, to train an image classifier to be robust to spurious correlations. We design the framework to iteratively detect and mitigate the spurious correlations that the classifier excessively relies on for predictions. To achieve this, we first propose to utilize a pre-trained vision-language model to extract text-format attributes from images. These attributes enable us to curate data with various class-attribute correlations, and we formulate a novel metric to measure the degree of these correlations' spuriousness. Then, to mitigate the reliance on spurious correlations, we propose a meta-learning strategy in which the support (training) sets and query (test) sets in tasks are curated with different spurious correlations that have high degrees of spuriousness. By meta-training the classifier on these spuriousness-aware meta-learning tasks, our classifier can learn to be invariant to the spurious correlations. We demonstrate that our method is robust to spurious correlations without knowing them a priori and achieves the best on five benchmark datasets with different robustness measures. Our code is available at https://github.com/gtzheng/SPUME. Guangtao Zheng, Wenqian Ye, Aidong Zhang 0001 |
KDD | 1 |
| 2023 | Learning to Learn Task Transformations for Improved Few-Shot ClassificationabstractMeta-learning has shown great promise in few-shot image classification where only a small amount of labeled data is available in each classification task. Many training tasks are provided to train a meta-model that can quickly learn new and similar concepts with few labeled samples. Data augmentation is often used to augment training tasks to avoid overfitting. However, existing data augmentation methods are often manually designed and fixed during training, ignoring training dynamics and the difference between various meta-learning settings specified by meta-model architectures and meta-learning algorithms. To address this problem, we add a task transformation layer between a training task and a meta-model such that the right amount of perturbation is added to training tasks for a certain meta-learning setting at a certain training stage. By jointly optimizing the task transformation layer and the meta-model, we avoid the risk of providing tasks that are either too easy or too difficult during training. We design the task transformation layer as a stochastic transformation function, adding the flexibility in how a training task can be transformed. We leverage differentiable data augmentations as the building blocks of the task transformation function for efficient optimization. Extensive experiments show that our method can consistently improve the few-shot generalization performance of various meta-models trained with different meta-learning algorithms, meta-model architectures, and datasets. Guangtao Zheng, Qiuling Suo, Mengdi Huai, Aidong Zhang 0001 |
SDM | 1 |
| 2022 | Knowledge-Guided Semantics Adjustment for Improved Few-Shot ClassificationabstractIn few-shot image classification, it is challenging for deep neural networks to infer the true class of an image when it contains multiple class-unrelated objects and its label carries no semantic meanings, such as Class 1 or Class 2. In contrast, knowing what are not important in a typical classification task, humans can quickly identify the right class objects with very few images. In this paper, we propose to extract semantic features from a given dataset to filter out class-unrelated objects in a few-shot task. Each semantic feature is meta-learned and represents a common object or pattern shared by many tasks. The strengths of these features in a given image are adjusted by an importance kernel encoding the meta-learned knowledge such that class-unrelated objects can be suppressed, and the few-shot classification performance can be improved. To facilitate learning and identifying semantic features in an image, we further propose an image representation decomposition module to decouple complex correlations between objects in an image embedding. The experimental analysis demonstrates the effectiveness of our method, especially in the extremely low-shot cases. Guangtao Zheng, Aidong Zhang 0001 |
ICDM | 1 |
| 2021 | Embeddings of genomic region sets capture rich biological associations in lower dimensionsabstractMOTIVATION: Genomic region sets summarize functional genomics data and define locations of interest in the genome such as regulatory regions or transcription factor binding sites. The number of publicly available region sets has increased dramatically, leading to challenges in data analysis. RESULTS: We propose a new method to represent genomic region sets as vectors, or embeddings, using an adapted word2vec approach. We compared our approach to two simpler methods based on interval unions or term frequency-inverse document frequency and evaluated the methods in three ways: First, by classifying the cell line, antibody or tissue type of the region set; second, by assessing whether similarity among embeddings can reflect simulated random perturbations of genomic regions; and third, by testing robustness of the proposed representations to different signal thresholds for calling peaks. Our word2vec-based region set embeddings reduce dimensionality from more than a hundred thousand to 100 without significant loss in classification performance. The vector representation could identify cell line, antibody and tissue type with over 90% accuracy. We also found that the vectors could quantitatively summarize simulated random perturbations to region sets and are more robust to subsampling the data derived from different peak calling thresholds. Our evaluations demonstrate that the vectors retain useful biological information in relatively lower-dimensional spaces. We propose that vector representation of region sets is a promising approach for efficient analysis of genomic region data. AVAILABILITY AND IMPLEMENTATION: https://github.com/databio/regionset-embedding. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Erfaneh Gharavi, Aaron Gu, Guangtao Zheng, Jason P. Smith, Hyun Jae Cho, Aidong Zhang 0001, Donald E. Brown, Nathan C. Sheffield |
Bioinform. | 3 |
| 2020 | Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionabstractGenerating explanations for neural networks has become crucial for their applications in real-world with respect to reliability and trustworthiness.In natural language processing, existing methods usually provide important features which are words or phrases selected from an input text as an explanation, but ignore the interactions between them.It poses challenges for humans to interpret an explanation and connect it to model prediction.In this work, we build hierarchical explanations by detecting feature interactions.Such explanations visualize how words and phrases are combined at different levels of the hierarchy, which can help users understand the decision-making of blackbox models.The proposed method is evaluated with three neural text classifiers (LSTM, CNN, and BERT) on two benchmark datasets, via both automatic and human evaluations.Experiments show the effectiveness of the proposed method in providing explanations that are both faithful to models and interpretable to humans. Guangtao Zheng, Yangfeng Ji |
ACL | 2 |
| 2019 | Multi-Layer Coding and Map-Assisted Partial Group Decoding for Multi-Color Multi-User Visible Light CommunicationabstractIn this paper, interference cancellation with multilayer coding and constrained partial group decoder (CPGD) is adopted for visible light communication (VLC) with a light emitting diode (LED) array. To reduce the online computational complexity, a decoding map on the decoding orders and the associated rate allocations is constructed via solving the max-min fairness problem based on the derived achievable rate of a VLC multiple access channel. A classification-based algorithm is proposed to reduce the decoding map size. Finally, the system throughput is further enhanced via solving the transmitter-user association problem followed by an iterative update of rate allocation and decoding order. From the numerical results, symmetric decoding map is observed for the squared transmitting LED array. Guangtao Zheng, Chen Gong 0001, Zhengyuan Xu |
ICC | 1 |
| 2019 | Constrained Partial Group Decoding With Max-Min Fairness for Multi-Color Multi-User Visible Light CommunicationabstractConsider a fixed orientation multi-user visible light communication (VLC) system with multi-color light emitting diodes (LEDs) where each transmitter serves a user with limited coordination impeding transmitter-side beamforming. In this paper, a multi-layer coding and constrained partial group decoding (CPGD) method is proposed to tackle strong color interference and increase the system throughput. After channel model formulation, user information rates are allocated and decoding order for all the received data layers is obtained by solving a max-min fairness problem using a greedy algorithm. An achievable rate is derived under the truncated Gaussian input distribution. To reduce the decoding complexity, a map on the decoding order and rate allocation is constructed for all positions of interest on the receiver plane parallel to the transmitter plane, and its size is reduced by a clustering algorithm. Meanwhile, the symmetrical geometry of LED arrays is exploited. Finally, the transmitter-user association problem is formulated and solved by a genetic algorithm. It is observed that the system throughput increases as the receivers are slightly misaligned with corresponding LED arrays due to the reduced interference level, but decreases afterwards due to the weakened link gain. Guangtao Zheng, Chen Gong 0001, Zhengyuan Xu |
IEEE Trans. Commun. | 1 |