EDBT 2026 Demo / reviewers in the wild / expert
Qinying Gu
dblp:365/4635
· DBLP profile ↗
17ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0002-8863-0810ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CFSM: A Novel Causal Feature Selection Module for Two-Dimensional Out-of-Distribution GeneralizationabstractIn real-world scenarios, training and test data are often collected in diverse settings, leading to domain shifts arising from evolving environments and selection bias. While causality-inspired methods have shown promising results in tackling the out-of-distribution (OOD) generalization issue, prior methods treat the discovered differences across domains as confounding variables. While effective in handling domain differences (i.e., unseen environmental features in test data), they may fail when confronted with intricate spurious correlations in real-world datasets. In this study, we first analyze this limitation to inadequate modeling of causal intervention and derive the OOD generalization bound to explain the challenges it introduces. To address this problem, we propose a modified causal intervention approach to mitigate various types of confounders. Motivated by the mathematical formulation of our modified causal intervention, we introduce the Causal Feature Selection Module (CFSM) to suppress model weights on both domain-differences features and spurious correlation features. Integrated within the Base Feature Extraction Module, In-Sample Module, and Cross-Sample Module (B-I-C architecture), CFSM collectively neutralizes the confounding effects arising from both domain discrepancies and correlation distinctions, thereby achieving causal feature selection. Under mild assumptions, we prove that the proposed CFSM method can achieve strictly lower OOD errors. Further experiments conducted on various benchmark datasets demonstrate the effectiveness of the proposed method. Compared to previous deconfounding methods, our method not only mitigates the effect of domain-differences features but also the hard-to-identify spurious correlation features, achieving significant improvements in two-dimensional OOD generalization. Weihan Yin, Yiyao Yang, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Visual Position Prompt for MLLM Based Visual GroundingabstractAlthough Multimodal Large Language Models (MLLMs) excel at various image-related tasks, they encounter challenges in precisely aligning coordinates with spatial information within images, particularly in position-aware tasks such as visual grounding. This limitation arises from two key factors. First, MLLMs lack explicit spatial references, making it difficult to associate textual descriptions with precise image locations. Second, their feature extraction processes prioritize global context over fine-grained spatial details, leading to weak localization capability. To address these issues, we introduce VPP-LLaVA, an MLLM enhanced with Visual Position Prompt (VPP) to improve its grounding capability. VPP-LLaVA integrates two complementary mechanisms: the global VPP overlays a learnable, axis-like tensor onto the input image to provide structured spatial cues, while the local VPP incorporates position-aware queries to support fine-grained localization. To effectively train our model with spatial guidance, we further introduce VPP-SFT, a curated dataset of 0.6 M high-quality visual grounding samples. Designed in a compact format, it enables efficient training and is significantly smaller than datasets used by other MLLMs (e.g., 21 M samples in MiniGPT-v2), yet still provides a strong performance boost. The resulting model, VPP-LLaVA, not only achieves state-of-the-art results on standard visual grounding benchmarks but also demonstrates strong zero-shot generalization to challenging unseen datasets. Wei Tang 0011, Yanpeng Sun, Qinying Gu, Zechao Li |
IEEE Trans. Multim. | 3 |
| 2025 | Enhancing Nursing and Elderly Care with Large Language Models: An AI-Driven FrameworkabstractThis paper explores the application of large language models (LLMs) in nursing and elderly care, focusing on AI-driven patient monitoring and interaction. We introduce a novel Chinese nursing dataset and implement incremental pre-training (IPT) and supervised fine-tuning (SFT) techniques to enhance LLM performance in specialized tasks. Using LangChain, we develop an interactable nursing assistant capable of real-time care and personalized interventions. Experimental results demonstrate significant improvements, paving the way for AI-driven solutions to meet the growing demands of healthcare in aging populations. Qiao Sun 0003, Jiexin Xie, Nanyang Ye 0001, Qinying Gu, Shijie Guo |
COLING | 4 |
| 2025 | Decision SpikeFormer: Spike-Driven Transformer for Decision MakingabstractOffline reinforcement learning (RL) enables policy training solely on pre-collected data, avoiding direct environment interaction—a crucial benefit for energy-constrained embodied AI applications. Although Artificial Neural Networks (ANN)-based methods perform well in offline RL, their high computational and energy demands motivate exploration of more efficient alternatives. Spiking Neural Networks (SNNs) show promise for such tasks, given their low power consumption. In this work, we introduce DSFormer, the first spike-driven transformer model designed to tackle offline RL via sequence modeling. Unlike existing SNN transformers focused on spatial dimensions for vision tasks, we develop Temporal Spiking Self-Attention (TSSA) and Positional Spiking Self-Attention (PSSA) in DSFormer to capture the temporal and positional dependencies essential for sequence modeling in RL. Additionally, we propose Progressive Threshold-dependent Batch Normalization (PTBN), which combines the benefits of LayerNorm and BatchNorm to preserve temporal dependencies while maintaining the spiking nature of SNNs. Comprehensive results in the D4RL benchmark show DSFormer’s superiority over both SNN and ANN counterparts, achieving 78.4% energy savings, highlighting DSFormer’s advantages not only in energy efficiency but also in competitive performance. Code and models are public at project page. Qinying Gu, Nanyang Ye 0001 |
CVPR | 2 |
| 2025 | OODD: Test-time Out-of-Distribution Detection with Dynamic DictionaryabstractOut-of-distribution (OOD) detection remains challenging for deep learning models, particularly when test-time OOD samples differ significantly from training outliers. We propose OODD, a novel test-time OOD detection method that dynamically maintains and updates an OOD dictionary without fine-tuning. Our approach leverages a priority queue-based dictionary that accumulates representative OOD features during testing, combined with an informative inlier sampling strategy for in-distribution (ID) samples. To ensure stable performance during early testing, we propose a dual OOD stabilization mechanism that leverages strategically generated outliers derived from ID data. To our best knowledge, extensive experiments on the OpenOOD benchmark demonstrate that OODD significantly outperforms existing methods, achieving a 26.0% improvement in FPR95 on CIFAR-100 Far OOD detection compared to the state-of-the-art approach. Furthermore, we present an optimized variant of the KNN-based OOD detection framework that achieves a 3x speedup while maintaining detection performance. Our code is available at https://github.com/zxk1212/OODD. Zewen Sun, Hengyu Liu 0007, Qinying Gu, Nanyang Ye 0001 |
CVPR | 5 |
| 2025 | Conformalized Causal Learning for Uncertainty-Aware Mineral Prospectivity Mapping
Evelyn Jessica Jaya, Qinying Gu, Xinbing Wang, Nanyang Ye 0001 |
ICANN (4) | 2 |
| 2025 | Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion ModelsabstractGiven a style-reference image as the additional image condition, text-to-image diffusion models have demonstrated impressive capabilities in generating images that possess the content of text prompts while adopting the visual style of the reference image. However, current state-of-the-art methods often struggle to disentangle content and style from style-reference images, leading to issues such as content leakages. To address this issue, we propose a masking-based method that efficiently decouples content from style without the need of tuning any model parameters. By simply masking specific elements in the style reference's image features, we uncover a critical yet under-explored principle: guiding with appropriately-selected fewer conditions (e.g., dropping several image feature elements) can efficiently avoid unwanted content flowing into the diffusion models, enhancing the style transfer performances of text-to-image diffusion models. In this paper, we validate this finding both theoretically and experimentally. Extensive experiments across various styles demonstrate the effectiveness of our masking-based method and support our theoretical results. Xinbing Wang, Chenghu Zhou, Qinying Gu, Nanyang Ye 0001 |
ICLR | 4 |
| 2025 | Generalizable Multi-Camera 3D Object Detection from a Single Source via Fourier Cross-View LearningabstractImproving the generalization of multi-camera 3D object detection is essential for safe autonomous driving in the real world. In this paper, we consider a realistic yet more challenging scenario, which aims to improve the generalization when only single source data available for training, as gathering diverse domains of data and collecting annotations is time-consuming and labor-intensive. To this end, we propose the Fourier Cross-View Learning (FCVL) framework including Fourier Hierarchical Augmentation (FHiAug), an augmentation strategy in the frequency domain to boost domain diversity, and Fourier Cross-View Semantic Consistency Loss to facilitate the model to learn more domain-invariant features from adjacent perspectives. Furthermore, we provide theoretical guarantees via augmentation graph theory. To the best of our knowledge, this is the first study to explore generalizable multi-camera 3D object detection with a single source. Extensive experiments on various testing domains have demonstrated that our approach achieves the best performance across various domain generalization methods. Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
ICML | 2 |
| 2025 | Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningabstractOffline reinforcement learning (offline RL) is increasingly approached as a sequence modeling task, with methods leveraging advanced architectures like Transformers to capture trajectory dependencies. Despite significant progress, the mechanisms underlying their effectiveness and limitations remain insufficiently understood. We conduct a thorough analysis on the representative Decision Transformer (DT) model using an entropy analysis and identify the inconsistencies in state-action-reward ($\langle s, a, R \rangle$) distributions causing attention ``dispersal". To address this, we propose a hierarchical framework that decomposes sequence modeling into intra-step relational modeling—handled by a Token Merger that fuses each $\langle s, a, R \rangle$ triplet—and inter-step modeling—handled by a Token Mixer across timesteps. We investigate several Token Merger designs and validate their effectiveness across various offline RL methods.
Furthermore, our theoretical analysis and experimental results suggest that while Token Mixers are important, lightweight architecture can also achieve even better performance to more complex ones. We therefore propose a parameter-free Average Pooling Token Mixer, which, combined with a convolutional Token Merger, forms our final model, Decision HiFormer (DHi). DHi achieves a \textbf{73.6\%} improvement in inference speed and an \textbf{9.3\%} gain in policy performance on the D4RL benchmark compared to DT. DHi also generalizes well to real-world robotic manipulation tasks, offering both practical benefits and insights into sequence-based policy design for offline RL. Code and models are public at \href{https://wei-nijuan.github.io/DecisionHiFormer/}{project page}. Leiyu Wang, Heyue Li, Luoyi Fan, Nanyang Ye 0001, Qinying Gu |
NeurIPS | 8 |
| 2025 | Δ Energy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
Xinbing Wang, Qinying Gu, Nanyang Ye 0001 |
NeurIPS | 4 |
| 2025 | Bayes-CAL: Robust Cross-Modal Alignment by Bayesian Approach for Few-Shot OoD Generalization
Weihan Yin, Fan Wu 0006, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | InfoBound: A Provable Information-Bounds Inspired Framework for Both OoD Generalization and OoD DetectionabstractIn real-world scenarios, distribution shifts give rise to the importance of two problems: out-of-distribution (OoD) generalization, which focuses on models' generalization ability against covariate shifts (i.e., the changes of environments), and OoD detection, which aims to be aware of semantic shifts (i.e., test-time unseen classes). Real-world testing environments often involve a combination of both covariate and semantic shifts. While numerous methods have been proposed to address these critical issues, only a few works tackled them simultaneously. Moreover, prior works often improve one problem but sacrifice the other. To overcome these limitations, we delve into boosting OoD detection and OoD generalization from the perspective of information theory, which can be easily applied to existing models and different tasks. Building upon the theoretical bounds for mutual information and conditional entropy, we provide a unified approach, composed of Mutual Information Minimization (MI-Min) and Conditional Entropy Maximizing (CE-Max). Extensive experiments and comprehensive evaluations on multi-label image classification and object detection have demonstrated the superiority of our method. It successfully mitigates trade-offs between the two challenges compared to competitive baselines. Zichao Nie, Yuan Gao 0050, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Domain Invariant Learning for Gaussian Processes and Bayesian ExplorationabstractOut-of-distribution (OOD) generalization has long been a challenging problem that remains largely unsolved. Gaussian processes (GP), as popular probabilistic model classes, especially in the small data regime, presume strong OOD generalization abilities. Surprisingly, their OOD generalization abilities have been under-explored before compared with other lines of GP research. In this paper, we identify that GP is not free from the problem and propose a domain invariant learning algorithm for Gaussian processes (DIL-GP) with a min-max optimization on the likelihood. DIL-GP discovers the heterogeneity in the data and forces invariance across partitioned subsets of data. We further extend the DIL-GP to improve Bayesian optimization's adaptability on changing environments. Numerical experiments demonstrate the superiority of DIL-GP for predictions on several synthetic and real-world datasets. We further demonstrate the effectiveness of the DIL-GP Bayesian optimization method on a PID parameters tuning experiment for a quadrotor. The full version and source code are available at: https://github.com/Billzxl/DIL-GP. Xilong Zhao, Siyuan Bian, Yaoyun Zhang, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
AAAI | 5 |
| 2024 | MiniConGTS: A Near Ultimate Minimalist Contrastive Grid Tagging Scheme for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) aims to co-extract the sentiment triplets in a given corpus.Existing approaches within the pretraining-finetuning paradigm tend to either meticulously craft complex tagging schemes and classification heads, or incorporate external semantic augmentation to enhance performance.In this study, we, for the first time, re-evaluate the redundancy in tagging schemes and the internal enhancement in pretrained representations.We propose a method to improve and utilize pretrained representations by integrating a minimalist tagging scheme and a novel token-level contrastive learning strategy.The proposed approach demonstrates comparable or superior performance compared to stateof-the-art techniques while featuring a more compact design and reduced computational overhead.Additionally, we are the first to formally evaluate GPT-4's performance in fewshot learning and Chain-of-Thought scenarios for this task.The results demonstrate that the pretraining-finetuning paradigm remains highly effective even in the era of large language models. Qiao Sun 0003, Liujia Yang, Minghao Ma, Nanyang Ye 0001, Qinying Gu |
EMNLP | 5 |
| 2024 | CRoFT: Robust Fine-Tuning with Concurrent Optimization for OOD Generalization and Open-Set OOD DetectionabstractRecent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often involve further fine-tuning of VL-PTMs, which may distort their general knowledge and impair their ability to handle distribution shifts. In real-world scenarios, machine learning systems inevitably encounter both covariate shifts (e.g., changes in image styles) and semantic shifts (e.g., test-time unseen classes). This highlights the importance of enhancing out-of-distribution (OOD) generalization on covariate shifts and simultaneously detecting semantic-shifted unseen classes. Thus a critical but underexplored question arises: How to improve VL-PTMs’ generalization ability to closed-set OOD data, while effectively detecting open-set unseen classes during fine-tuning? In this paper, we propose a novel objective function of OOD detection that also serves to improve OOD generalization. We show that minimizing the gradient magnitude of energy scores on training data leads to domain-consistent Hessians of classification loss, a strong indicator for OOD generalization revealed by theoretical analysis. Based on this finding, we have developed a unified fine-tuning framework that allows for concurrent optimization of both tasks. Extensive experiments have demonstrated the superiority of our method. The code is available at https://github.com/LinLLLL/CRoFT. Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
ICML | 3 |
| 2024 | Vision-Language Alignment Learning Under Affinity and Divergence Principles for Few-Shot Out-of-Distribution Generalization
Weihan Yin, Yiyao Yang, Fan Wu 0006, Zhaoyu Zeng, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
Int. J. Comput. Vis. | 6 |
| 2024 | OoD-Control: Generalizing Control in Unseen EnvironmentsabstractGeneralizing out-of-distribution (OoD) is critical but challenging in real applications such as unmanned aerial vehicle (UAV) flight control. Previous machine learning-based control has shown promise in dealing with complex real-world environments but suffers huge performance degradation facing OoD scenarios, posing risks to the stability and safety of UAVs. In this paper, we found that the introduced random noises during training surprisingly yield theoretically guaranteed performances via a proposed functional optimization framework. More encouragingly, this framework does not involve common Lyapunov assumptions used in this field, making it more widely applicable. With this framework, the upperbound for control error is induced. We also proved that the induced random noises can lead to lower OoD control errors. Based on our theoretical analysis, we further propose OoD-Control to generalize control in unseen environments. Numerical experiments demonstrate the superiority of the proposed algorithm, surpassing previous state-of-the-art by 65% under challenging unseen environments. We further extend to outdoor real-world experiments and found that the control error is reduced by 50% approximately. Nanyang Ye 0001, Zhaoyu Zeng, Jundong Zhou, Yuxiao Duan, Haoqi Zeng, Qinying Gu, Xinbing Wang, Chenghu Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |