VLDB 2026 Research / reviewers in the wild / expert
Li Gu
dblp:87/782
· DBLP profile ↗
17ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ETR: Entropy Trend Reward for Efficient Chain-of-Thought ReasoningabstractChain-of-thought (CoT) reasoning improves large language model performance on complex tasks, but often produces excessively long and inefficient reasoning traces.Existing methods shorten CoTs using length penalties or global entropy reduction, implicitly assuming that low uncertainty is desirable throughout reasoning.We show instead that reasoning efficiency is governed by the trajectory of uncertainty.CoTs with dominant downward entropy trends are substantially shorter.Motivated by this insight, we propose Entropy Trend Reward (ETR), a trajectory-aware objective that encourages progressive uncertainty reduction while allowing limited local exploration.We integrate ETR into Group Relative Policy Optimization (GRPO) and evaluate it across multiple reasoning models and challenging benchmarks.ETR consistently achieves a superior accuracy-efficiency trade-off, improving DeepSeek-R1-Distill-7B by +9.9% accuracy while reducing CoT length by 67% across four benchmarks. Xuan Xiong, Huan Liu 0014, Li Gu, Zhixiang Chi, Yuanhao Yu, Yang Wang 0003 |
ACL (1) | 3 |
| 2025 | MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt TuningabstractRecent advancements in handwritten text recognition (HTR) have enabled the effective conversion of handwritten text to digital formats. However, achieving robust recognition across diverse writing styles remains challenging. Traditional HTR methods lack writer-specific personalization at test time due to limitations in model architecture and training strategies. Existing attempts to bridge this gap, through gradient-based meta-learning, still require labeled examples and suffer from parameter-inefficient fine-tuning, leading to substantial computational and memory overhead. To overcome these challenges, we propose an efficient framework that formulates personalization as prompt tuning, incorporating an auxiliary image reconstruction task with a self-supervised loss to guide prompt adaptation with unlabeled test-time examples. To ensure self-supervised loss effectively minimizes text recognition error, we leverage meta-learning to learn the optimal initialization of the prompts. As a result, our method allows the model to efficiently capture unique writing styles by updating less than 1% of its parameters and eliminating the need for time-intensive annotation processes. We validate our approach on the RIMES and IAM Handwriting Database benchmarks, where it consistently outperforms previous state-ofthe-art methods while using 20x fewer parameters. We believe this represents a significant advancement in personalized handwritten text recognition, paving the way for more reliable and practical deployment in resource-constrained scenarios. Wenhao Gu, Li Gu, Chingyee Yee Suen, Yang Wang 0003 |
CVPR | 2 |
| 2025 | Plug-in Feedback Self-Adaptive Attention in CLIP for Training-Free Open-Vocabulary SegmentationabstractCLIP exhibits strong visual-textual alignment but struggle with open-vocabulary segmentation due to poor localization. Prior methods enhance spatial coherence by modifying intermediate attention. But, this coherence isn't consistently propagated to the final output due to subsequent operations such as projections. Additionally, intermediate attention lacks direct interaction with text representations, such semantic discrepancy limits the full potential of CLIP. In this work, we propose a training-free, feedback-driven self-adaptive framework that adapts output-based patch-level correspondences back to the intermediate attention. The output predictions, being the culmination of the model's processing, encapsulate the most comprehensive visual and textual semantics about each patch. Our approach enhances semantic consistency between internal representations and final predictions by leveraging the model's outputs as a stronger spatial coherence prior. We design key modules, including attention isolation, confidence-based pruning for sparse adaptation, and adaptation ensemble, to effectively feedback the output coherence cues. Our method functions as a plug-in module, seamlessly integrating into four state-of-the-art approaches with three backbones (ViT-B, ViT-L, ViT-H). We further validate our framework across multiple attention types (Q-K, self-self, and Proxy augmented with MAE, SAM, and DINO). Our approach consistently improves their performance across eight benchmarks. Zhixiang Chi, Li Gu, Huan Liu 0014, Ziqiang Wang 0003, Yang Wang 0003, Konstantinos N. Plataniotis |
ICCV | 3 |
| 2025 | Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain AdaptationabstractFew-shot Test-Time Domain Adaptation focuses on adapting a model at test time to a specific domain using only a few unlabeled examples, addressing domain shift. Prior methods leverage CLIP's strong out-of-distribution (OOD) abilities by generating domain-specific prompts to guide its generalized, frozen features. However, since downstream datasets are not explicitly seen by CLIP, solely depending on the feature space knowledge is constrained by CLIP's prior knowledge. Notably, when using a less robust backbone like ViT-B/16, performance significantly drops on challenging real-world benchmarks. Departing from the state-of-the-art of inheriting the intrinsic OOD capability of CLIP, this work introduces learning directly on the input space to complement the dataset-specific knowledge for frozen CLIP. Specifically, an independent side branch is attached in parallel with CLIP and enforced to learn exclusive knowledge via revert attention. To better capture the dataset-specific label semantics for downstream adaptation, we propose to enhance the inter-dispersion among text features via greedy text ensemble and refinement. The text and visual features are then progressively fused in a domain-aware manner by a generated domain prompt to adapt toward a specific domain. Extensive experiments show our method's superiority on 5 large-scale benchmarks (WILDS and DomainNet), notably improving over smaller networks like ViT-B/16 with gains of \textbf{+5.1} in F1 for iWildCam and \textbf{+3.1\%} in WC Acc for FMoW. \href{https://github.com/chi-chi-zx/L2C}{Our Code: L2C} Zhixiang Chi, Li Gu, Huan Liu 0014, Ziqiang Wang 0003, Yang Wang 0003, Konstantinos N. Plataniotis |
ICLR | 2 |
| 2025 | PointMAC: Meta-Learned Adaptation for Robust Test-Time Point Cloud CompletionabstractPoint cloud completion is essential for robust 3D perception in safety-critical applications such as robotics and augmented reality. However, existing models perform static inference and rely heavily on inductive biases learned during training, limiting their ability to adapt to novel structural patterns and sensor-induced distortions at test time.
To address this limitation, we propose PointMAC, a meta-learned framework for robust test-time adaptation in point cloud completion. It enables sample-specific refinement without requiring additional supervision.
Our method optimizes the completion model under two self-supervised auxiliary objectives that simulate structural and sensor-level incompleteness.
A meta-auxiliary learning strategy based on Model-Agnostic Meta-Learning (MAML) ensures that adaptation driven by auxiliary objectives is consistently aligned with the primary completion task.
During inference, we adapt the shared encoder on-the-fly by optimizing auxiliary losses, with the decoder kept fixed. To further stabilize adaptation, we introduce Adaptive $\lambda$-Calibration, a meta-learned mechanism for balancing gradients between primary and auxiliary objectives.
Extensive experiments on synthetic, simulated, and real-world datasets demonstrate that PointMAC achieves state-of-the-art results by refining each sample individually to produce high-quality completions. To the best of our knowledge, this is the first work to apply meta-auxiliary test-time adaptation to point cloud completion. Linlian Jiang, Li Gu, Ziqiang Wang 0003, Xinxin Zuo, Yang Wang 0003 |
NeurIPS | 3 |
| 2025 | DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary LearningabstractDespite recent significant advancements in Handwritten Document Recognition (HDR), the efficient and accurate recognition of text against complex backgrounds, diverse handwriting styles, and varying document layouts remains a practical challenge. Moreover, this issue is seldom addressed in academic research, particularly in scenarios with minimal annotated data available. In this paper, we introduce the DocTTT framework to address these challenges. The key innovation of our approach is that it uses test-time training to adapt the model to each specific input during testing. We propose a novel Meta-Auxiliary learning approach that combines Meta-learning and self-supervised Masked Autoencoder (MAE). During testing, we adapt the visual representation parameters using a self-supervised MAE loss. During training, we learn the model parameters using a meta-learning framework, so that the model parameters are learned to adapt to a new input effectively. Experimental results show that our proposed method significantly outperforms existing state-of-the-art approaches on benchmark datasets. Wenhao Gu, Li Gu, Ziqiang Wang 0003, Ching Y. Suen, Yang Wang 0003 |
WACV | 2 |
| 2025 | MTCR: Method for Matching Texts Against Causal RelationshipabstractText matching is considered a vital task in natural language processing and constitutes a fundamental component of many NLP applications. In practical scenarios, there exist numerous texts with causal relationships, such as questions and answers in question-answering systems. For text matching tasks involving causal relationships, considering the causal connections between text pairs is of paramount importance. Furthermore, when dealing with lengthy texts, causal signals often exhibit sparsity, rendering the extraction of causal features a challenging endeavor. To address the aforementioned issues, this paper introduces an approach that amalgamates causal knowledge distillation and causal semantic extraction, denoted as the Method for Matching Texts with Causal Relationship (MTCR). This framework effectively learns deep semantic representations of causal relationships between texts. MTCR excels at handling text matching tasks that involve causal relationships, including tasks like natural language inference and answer selection. Simultaneously, it effectively identifies instances of causal inversions. Through experimentation on five benchmark text matching datasets, our research findings indicate that the proposed method can effectively handle text matching tasks involve causal relationships. XinYue Jiang, Jingsong He, Li Gu |
Neural Process. Lett. | 3 |
| 2024 | Distribution Alignment for Fully Test-Time Adaptation with Dynamic Online Data Streams
Ziqiang Wang 0003, Zhixiang Chi, Li Gu, Zhi Liu 0003, Konstantinos N. Plataniotis, Yang Wang 0003 |
ECCV (24) | 4 |
| 2024 | Adapting to Distribution Shift by Visual Domain Prompt GenerationabstractIn this paper, we aim to adapt a model at test-time using a few unlabeled data to address distribution shifts.
To tackle the challenges of extracting domain knowledge from a limited amount of data, it is crucial to utilize correlated information from pre-trained backbones and source domains. Previous studies fail to utilize recent foundation models with strong out-of-distribution generalization. Additionally, domain-centric designs are not flavored in their works. Furthermore, they employ the process of modelling source domains and the process of learning to adapt independently into disjoint training stages. In this work, we propose an approach on top of the pre-computed features of the foundation model. Specifically, we build a knowledge bank to learn the transferable knowledge from source domains. Conditioned on few-shot target data, we introduce a domain prompt generator to condense the knowledge bank into a domain-specific prompt. The domain prompt then directs the visual features towards a particular domain via a guidance module. Moreover, we propose a domain-aware contrastive loss and employ meta-learning to facilitate domain knowledge extraction. Extensive experiments are conducted to validate the domain knowledge extraction. The proposed method outperforms previous work on 5 large-scale benchmarks including WILDS and DomainNet. Zhixiang Chi, Li Gu, Tao Zhong 0003, Huan Liu 0014, Yuanhao Yu, Konstantinos N. Plataniotis, Yang Wang 0003 |
ICLR | 2 |
| 2024 | Crowd Counting Using Meta-Test-Time AdaptationabstractMachine learning algorithms are commonly used for quickly and efficiently counting people from a crowd. Test-time adaptation methods for crowd counting adjust model parameters and employ additional data augmentation to better adapt the model to the specific conditions encountered during testing. The majority of current studies concentrate on unsupervised domain adaptation. These approaches commonly perform hundreds of epochs of training iterations, requiring a sizable number of unannotated data of every new target domain apart from annotated data of the source domain. Unlike these methods, we propose a meta-test-time adaptive crowd counting approach called CrowdTTA, which integrates the concept of test-time adaptation into the meta-learning framework and makes it easier for the counting model to adapt to the unknown test distributions. To facilitate the reliable supervision signal at the pixel level, we introduce uncertainty by inserting the dropout layer into the counting model. The uncertainty is then used to generate valuable pseudo labels, serving as effective supervisory signals for adapting the model. In the context of meta-learning, one image can be regarded as one task for crowd counting. In each iteration, our approach is a dual-level optimization process. In the inner update, we employ a self-supervised consistency loss function to optimize the model so as to simulate the parameters update process that occurs during the test phase. In the outer update, we authentically update the parameters based on the image with ground truth, improving the model's performance and making the pseudo labels more accurate in the next iteration. At test time, the input image is used for adapting the model before testing the image. In comparison to various supervised learning and domain adaptation methods, our results via extensive experiments on diverse datasets showcase the general adaptive capability of our approach across datasets with varying crowd densities and scales. Ferrante Neri, Li Gu, Ziqiang Wang 0003, Jian Wang 0110, Anyong Qing, Yang Wang 0003 |
Int. J. Neural Syst. | 3 |
| 2022 | MetaFSCIL: A Meta-Learning Approach for Few-Shot Class Incremental LearningabstractIn this paper, we tackle the problem of few-shot class incremental learning (FSCIL). FSCIL aims to incrementally learn new classes with only a few samples in each class. Most existing methods only consider the incremental steps at test time. The learning objective of these methods is often hand-engineered and is not directly tied to the objective (i.e. incrementally learning new classes) during testing. Those methods are sub-optimal due to the misalignment between the training objectives and what the methods are expected to do during evaluation. In this work, we proposed a bi-level optimization based on meta-learning to directly optimize the network to learn how to incrementally learn in the setting of FSCIL. Concretely, we propose to sample sequences of incremental tasks from base classes for training to simulate the evaluation protocol. For each task, the model is learned using a meta-objective such that it is capable to perform fast adaptation without forgetting. Furthermore, we propose a bi-directional guided modulation, which is learned to automatically modulate the activations to reduce catastrophic forgetting. Extensive experimental results demonstrate that the proposed method outperforms the baseline and achieves the state-of-the-art results on CIFARIOO, MiniImageNet, and CUB200 datasets. Zhixiang Chi, Li Gu, Huan Liu 0014, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005 |
CVPR | 2 |
| 2022 | Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay
Huan Liu 0014, Li Gu, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005 |
ECCV (24) | 2 |
| 2022 | Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-ExpertsabstractIn this paper, we tackle the problem of domain shift. Most existing methods perform training on multiple source domains using a single model, and the same trained model is used on all unseen target domains. Such solutions are sub-optimal as each target domain exhibits its own specialty, which is not adapted. Furthermore, expecting single-model training to learn extensive knowledge from multiple source domains is counterintuitive. The model is more biased toward learning only domain-invariant features and may result in negative knowledge transfer. In this work, we propose a novel framework for unsupervised test-time adaptation, which is formulated as a knowledge distillation process to address domain shift. Specifically, we incorporate Mixture-of-Experts (MoE) as teachers, where each expert is separately trained on different source domains to maximize their specialty. Given a test-time target domain, a small set of unlabeled data is sampled to query the knowledge from MoE. As the source domains are correlated to the target domains, a transformer-based aggregator then combines the domain knowledge by examining the interconnection among them. The output is treated as a supervision signal to adapt a student prediction network toward the target domain. We further employ meta-learning to enforce the aggregator to distill positive knowledge and the student network to achieve fast adaptation. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art and validates the effectiveness of each proposed component. Our code is available at https://github.com/n3il666/Meta-DMoE. Tao Zhong 0003, Zhixiang Chi, Li Gu, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005 |
NeurIPS | 3 |
| 2020 | Exploring Spatial-Temporal Representations for fNIRS-based Intimacy Detection via an Attention-enhanced Cascade Convolutional Recurrent Neural NetworkabstractThe detection of intimacy plays a crucial role in the improvement of intimate relationship, which contributes to promote the family and social harmony. Previous studies have shown that different degrees of intimacy have significant differences in brain imaging. Recently, work has emerged to recognise intimacy automatically by using machine learning techniques. Moreover, considering the temporal dynamic characteristics of intimacy relationship on neural mechanism, how to model spatiotemporal dynamics for intimacy prediction effectively is still a challenge. In this paper, we propose a novel method to explore deep spatial-temporal representations for intimacy prediction by anAttention-enhancedCascadeConvolutionalRecurrentNeuralNetwork(ACCRNN). Given the advantages of time-frequency resolution in complex neuronal activities analysis, this paper utilizesfunctionalnear-infraredspectroscopy(fNIRS) to analyse and infer intimate relationship. We collected fNIRS-based dataset for the analysis of intimate relationship. Forty-two-channel fNIRS signals are recorded from the 44 subjects' prefrontal cortex when they watched a total of 18 photos of lovers, friends and strangers for 30 seconds per photo. The experimental results show that our proposed method outperforms the others in terms of accuracy with the precision of 96.5%. To the best of our knowledge, this is the first time that such a hybrid deep architecture has been employed for fNIRS-based intimacy prediction. Ziping Zhao 0001, Li Gu, Björn W. Schuller |
ICPR | 4 |
| 2019 | DMM-Net: Differentiable Mask-Matching Network for Video Object SegmentationabstractIn this paper, we propose the differentiable mask-matching network (DMM-Net) for solving the video object segmentation problem where the initial object masks are provided. Relying on the Mask R-CNN backbone, we extract mask proposals per frame and formulate the matching between object templates and proposals as a linear assignment problem where thA heading inside a blocke cost matrix is predicted by a deep convolutional neural network. We propose a differentiable matching layer which unrolls a projected gradient descent algorithm in which the projection step exploits the Dykstra's algorithm. We prove that under mild conditions, the matching is guaranteed to converge to the optimal one. In practice, it achieves similar performance compared to the Hungarian algorithm during inference. Meanwhile, we can back-propagate through it to learn the cost matrix. After matching, a U-Net style architecture is exploited to refine the matched mask per time step. On DAVIS 2017 dataset, DMM-Net achieves the best performance without online learning on the first frames and the 2nd best with it. Without any fine-tuning, DMM-Net performs comparably to state-of-the-art methods on SegTrack v2 dataset. At last, our differentiable matching layer is very simple to implement; we attach the PyTorch code in the supplementary material which is less than 50 lines long. Xiaohui Zeng, Renjie Liao 0001, Li Gu, Yuwen Xiong, Sanja Fidler, Raquel Urtasun |
ICCV | 3 |
| 2019 | Analysing and Inferring of Intimacy Based on fNIRS Signals and Peripheral Physiological SignalsabstractIntimacy refers to a relatively long-lasting affinity relationship between individuals, which involves complex neuronal activities and physiological changes in the body. Recent advancements in the field of neuroimaging have demonstrated that functional near-infrared spectroscopy (fNIRS) has excellent potential for intimate relationship analysis. Signals such as fNIRS and physiological signals are increasingly utilised in this regard due to their consistency and complementarity. In this paper, first, we apply fNIRS and physiological database collected from 26 subjects when viewing lover, friend and stranger pictures to analyse and infer the intimacy. Then, the time domain information from both the fNIRS and physiological signals are utilised to exploit the representation of intimacy by General Linear Model (GLM) and Complex Brain Network Analysis (CBNA) methods. Based on these two methods, the intimacy can be analysed with different brain activation patterns. Finally, different machine learning techniques are utilised to predict the intimate relationship. The results demonstrate that multi-modal features are more efficient for intimacy research. Moreover, the average classification accuracy of ensemble learning is 98.72% whereas for KNN it is 91.03%. Ziping Zhao 0001, Li Gu, Nicholas Cummins, Björn W. Schuller |
IJCNN | 4 |
| 2018 | Adversarial Distillation of Bayesian Neural Network PosteriorsabstractBayesian neural networks (BNNs) allow us to reason about uncertainty in a principled way. Stochastic Gradient Langevin Dynamics (SGLD) enables efficient BNN learning by drawing samples from the BNN posterior using mini-batches. However, SGLD and its extensions require storage of many copies of the model parameters, a potentially prohibitive cost, especially for large neural networks. We propose a framework, Adversarial Posterior Distillation, to distill the SGLD samples using a Generative Adversarial Network (GAN). At test-time, samples are generated by the GAN. We show that this distillation framework incurs no loss in performance on recent BNN applications including anomaly detection, active learning, and defense against adversarial attacks. By construction, our framework distills not only the Bayesian predictive distribution, but the posterior itself. This allows one to compute quantities such as the approximate model variance, which is useful in downstream tasks. To our knowledge, these are the first results applying MCMC-based BNNs to the aforementioned applications. Kuan-Chieh Wang, Paul Vicol, James Lucas, Li Gu, Roger B. Grosse, Richard S. Zemel |
ICML | 4 |