EDBT 2026 Demo / reviewers in the wild / expert
Yi Zhang 0109
dblp:64/6544-109
· DBLP profile ↗
12ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-5831-0170ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TLoRA: Task-aware Low Rank Adaptation of Large Language ModelsabstractLow-Rank Adaptation (LoRA) has become a widely adopted parameter-efficient fine-tuning method for large language models, with its effectiveness largely influenced by the allocation of ranks and scaling factors, as well as initialization.Existing LoRA variants typically address only one of these factors, often at the cost of increased training complexity or reduced practical efficiency.In this work, we present Task-aware Low-Rank Adaptation (TLoRA), a unified framework that jointly optimizes initialization and resource allocation at the outset of training.TLoRA introduces a data-driven initialization strategy that aligns the LoRA A matrix with task-relevant subspaces by performing singular value decomposition on the product of pre-trained weights and input activation covariance.After this, the A matrix is frozen, and only the B matrix is trained.Furthermore, TLoRA employs a sensitivitybased importance metric to adaptively allocate ranks and scaling factors across layers under a fixed parameter budget.We conduct extensive experiments that demonstrate TLoRA consistently performs excellently across various tasks, including natural language understanding, commonsense reasoning, math reasoning, code generation, and chat generation, while significantly reducing the number of trainable parameters. Weicheng Lin, Yi Zhang 0109, Jiawei Dang, Liang-Jie Zhang |
ACL (1) | 2 |
| 2026 | Training-Free Dual Hyperbolic Adapters for Better Cross-Modal ReasoningabstractRecent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational resources for fine-tuning in new domains. To address this issue, we develop a new adaptation method for large vision-language models, calledTraining-free Dual Hyperbolic Adapters(T-DHA). We characterize vision-language relationship between semantic concepts, which typically has a hierarchical tree structure, in the hyperbolic space instead of the traditional Euclidean space. Hyperbolic spaces exhibit exponential volume growth with radius, unlike the polynomial growth in Euclidean space. We find that this unique property is particularly effective for embedding hierarchical data structures using the Poincaré ball model, achieving significantly improved representation and discrimination power. Coupled with negative learning, it provides more accurate and robust classifications with fewer feature dimensions. Our extensive experimental results on various datasets demonstrate that the T-DHA method significantly outperforms existing state-of-the-art methods in few-shot image recognition and domain generalization tasks. Yi Zhang 0109, Chun-Wun Cheng, Ke Yu 0004, Yushun Tang, Carola-Bibiane Schönlieb, Zhihai He, Angelica I. Avilés-Rivero |
IEEE Trans. Multim. | 1 |
| 2025 | Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential EquationsabstractWe introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP’s robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance. Yi Zhang 0109, Chun-Wun Cheng, Zhihai He, Carola-Bibiane Schönlieb, Yuyan Chen, Angelica I. Avilés-Rivero |
AAAI | 1 |
| 2024 | Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsabstractContrastive Language-Image Pretraining (CLIP) model has exhibited remarkable efficacy in establishing cross-modal connections between texts and images, yielding impressive performance across a broad spectrum of downstream applications through fine-tuning. However, for generalization tasks, the current fine-tuning methods for CLIP, such as CoOp and CoCoOp, demonstrate relatively low performance on some fine-grained datasets. We recognize the underlying reason is that these previous methods only projected global features into the prompt, neglecting the various visual concepts, such as colors, shapes, and sizes, which are naturally transferable across domains and play a crucial role in generalization tasks. To address this issue, in this work, we propose Concept-Guided Prompt Learning (CPL) for vision-language models. Specifically, we leverage the well-learned knowledge of CLIP to create a visual concept cache to enable conceptguided prompting. In order to refine the text features, we further develop a projector that transforms multi-level visual features into text features. We observe that this concept-guided prompt learning approach is able to achieve enhanced consistency between visual and linguistic modalities. Extensive experimental results demonstrate that our CPL method significantly improves generalization capabilities compared to the current state-of-the-art methods. Yi Zhang 0109, Ce Zhang 0009, Ke Yu 0004, Yushun Tang, Zhihai He |
AAAI | 1 |
| 2024 | Conceptual Codebook Learning for Vision-Language Models
Yi Zhang 0109, Ke Yu 0004, Zhihai He |
ECCV (77) | 1 |
| 2024 | Test-Time Distribution Learning Adapter for Cross-Modal Visual ReasoningabstractVision-Language Pre-Trained (VLP) models, such as CLIP, have demonstrated remarkable effectiveness in learning generic visual representations. Several approaches aim to efficiently adapt VLP models to downstream tasks with limited supervision, aiming to leverage the acquired knowledge from VLP models. However, these methods suffer from either introducing biased representations or requiring high computational complexity, which hinders their effectiveness in fine-tuning the CLIP model. Moreover, when a model is trained on data specific to a particular domain, its ability to generalize to uncharted domains diminishes. In this work, we propose Test-Time Distribution LearNing Adapter (TT-DNA) which directly works during the testing period. Specifically, we estimate Gaussian distributions to model visual features of the few-shot support images to capture the knowledge from the support set. The cosine similarity between query image and the feature distribution of support images is used as the prediction of visual adapter. Subsequently, the visual adapter’s prediction merges with the original CLIP prediction via a residual connection, resulting in the final prediction. Our extensive experimental results on visual reasoning for human object interaction demonstrate that our proposed TT-DNA outperforms existing state-of-the-art methods by large margins. Yi Zhang 0109, Ce Zhang 0009 |
ICASSP | 1 |
| 2024 | Domain-Conditioned Transformer for Fully Test-time Adaptation
Yushun Tang, Shuoshuo Chen, Jiyuan Jia, Yi Zhang 0109, Zhihai He |
ACM Multimedia | 4 |
| 2024 | Training-Free Feature Reconstruction with Sparse Optimization for Vision-Language ModelsabstractIn this paper, we address the challenge of adapting vision-language models (VLMs) to few-shot image recognition in a training-free manner. We observe that existing methods are not able to effectively characterize the semantic relationship between support and query samples in a training-free setting. We recognize that, in the semantic feature space, the feature of the query image is a linear and sparse combination of support image features since support-query pairs are from the class and share the same small set of distinctive visual attributes. Motivated by this interesting observation, we propose a novel method called Training-free Feature ReConstruction with Sparse optimization (TaCo), which formulates the few-shot image recognition task as a feature reconstruction and sparse optimization problem. Specifically, we exploit the VLM to encode the query and support images into features. We utilize sparse optimization to reconstruct the query feature from the corresponding support features. The feature reconstruction error is then used to define the reconstruction similarity. Coupled with the text-image similarity provided by the VLM, our reconstruction similarity analysis accurately characterizes the relationship between support and query images. This results in significantly improved performance in few-shot image recognition. Our extensive experimental results on few-shot recognition demonstrate that our method outperforms existing state-of-the-art approaches by substantial margins. Yi Zhang 0109, Ke Yu 0004, Angelica I. Avilés-Rivero, Jiyuan Jia, Yushun Tang, Zhihai He |
ACM Multimedia | 1 |
| 2024 | Learning to Adapt CLIP for Few-Shot Monocular Depth EstimationabstractPre-trained Vision-Language Models (VLMs), such as CLIP, have shown enhanced performance across a range of tasks that involve the integration of visual and linguistic modalities. When CLIP is used for depth estimation tasks, the patches, divided from the input images, can be combined with a series of semantic descriptions of the depth information to obtain similarity results. The coarse estimation of depth is then achieved by weighting and summing the depth values, called depth bins, corresponding to the predefined semantic descriptions. The zero-shot approach circumvents the computational and time-intensive nature of traditional fully-supervised depth estimation methods. However, this method, utilizing fixed depth bins, may not effectively generalize as images from different scenes may exhibit distinct depth distributions. To address this challenge, we propose a few-shot-based method which learns to adapt the VLMs for monocular depth estimation to balance training costs and generalization capabilities. Specifically, it assigns different depth bins for different scenes, which can be selected by the model during inference. Additionally, we incorporate learnable prompts to preprocess the input text to convert the easily human-understood text into easily model-understood vectors and further enhance the performance. With only one image per scene for training, our extensive experiment results on the NYU V2 and KITTI dataset demonstrate that our method outperforms the previous state-of-the-art method by up to 10.6% in terms of MARE1. Xueting Hu, Ce Zhang 0009, Yi Zhang 0109, Bowen Hai, Ke Yu 0004, Zhihai He |
WACV | 3 |
| 2024 | Cross-Modal Concept Learning and Inference for Vision-Language Models
Yi Zhang 0109, Ce Zhang 0009, Yushun Tang, Zhihai He |
Neurocomputing | 1 |
| 2023 | BDC-Adapter: Brownian Distance Covariance for Better Vision-Language Reasoning
Yi Zhang 0109, Ce Zhang 0009, Yushun Tang, Zhihai He |
BMVC | 1 |
| 2023 | Unsupervised Prototype Adapter for Vision-Language Models
Yi Zhang 0109, Ce Zhang 0009, Xueting Hu, Zhihai He |
PRCV (1) | 1 |