EDBT 2026 Demo / reviewers in the wild / expert
Yushun Tang
dblp:341/5513
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-8350-7637ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Dual Hyperbolic Adapters for Better Cross-Modal ReasoningabstractRecent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational resources for fine-tuning in new domains. To address this issue, we develop a new adaptation method for large vision-language models, calledTraining-free Dual Hyperbolic Adapters(T-DHA). We characterize vision-language relationship between semantic concepts, which typically has a hierarchical tree structure, in the hyperbolic space instead of the traditional Euclidean space. Hyperbolic spaces exhibit exponential volume growth with radius, unlike the polynomial growth in Euclidean space. We find that this unique property is particularly effective for embedding hierarchical data structures using the Poincaré ball model, achieving significantly improved representation and discrimination power. Coupled with negative learning, it provides more accurate and robust classifications with fewer feature dimensions. Our extensive experimental results on various datasets demonstrate that the T-DHA method significantly outperforms existing state-of-the-art methods in few-shot image recognition and domain generalization tasks. Yi Zhang 0109, Chun-Wun Cheng, Ke Yu 0004, Yushun Tang, Carola-Bibiane Schönlieb, Zhihai He, Angelica I. Avilés-Rivero |
IEEE Trans. Multim. | 5 |
| 2025 | Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time ControlabstractReinforcement learning (RL) has achieved significant success across a wide range of domains, however, most existing methods are formulated in discrete time. In this work, we introduce a novel RL method for continuous-time control, where stochastic differential equations govern state-action dynamics. Departing from traditional value function-based approaches, our key contribution is the characterization of continuous-time Q-functions via a martingale condition and the linking of diffusion policy scores to the action gradient of a learned continuous Q-function by the dynamic programming principle. This insight motivates Continuous Q-Score Matching (CQSM), a score-based policy improvement algorithm. Notably, our method addresses a long-standing challenge in continuous-time RL: preserving the action-evaluation capability of Q-functions without relying on time discretization. We further provide theoretical closed-form solutions for linear-quadratic (LQ) control problems within our framework. Numerical results in simulated environments demonstrate the effectiveness of our proposed method and compare it to popular baselines. Chengxiu Hua, Jiawen Gu, Yushun Tang |
NeurIPS | 3 |
| 2025 | LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
Weiming Chen 0001, Yushun Tang, Zhihai He |
PRCV (6) | 3 |
| 2024 | Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsabstractContrastive Language-Image Pretraining (CLIP) model has exhibited remarkable efficacy in establishing cross-modal connections between texts and images, yielding impressive performance across a broad spectrum of downstream applications through fine-tuning. However, for generalization tasks, the current fine-tuning methods for CLIP, such as CoOp and CoCoOp, demonstrate relatively low performance on some fine-grained datasets. We recognize the underlying reason is that these previous methods only projected global features into the prompt, neglecting the various visual concepts, such as colors, shapes, and sizes, which are naturally transferable across domains and play a crucial role in generalization tasks. To address this issue, in this work, we propose Concept-Guided Prompt Learning (CPL) for vision-language models. Specifically, we leverage the well-learned knowledge of CLIP to create a visual concept cache to enable conceptguided prompting. In order to refine the text features, we further develop a projector that transforms multi-level visual features into text features. We observe that this concept-guided prompt learning approach is able to achieve enhanced consistency between visual and linguistic modalities. Extensive experimental results demonstrate that our CPL method significantly improves generalization capabilities compared to the current state-of-the-art methods. Yi Zhang 0109, Ce Zhang 0009, Ke Yu 0004, Yushun Tang, Zhihai He |
AAAI | 4 |
| 2024 | Window-Based Channel Attention for Wavelet-Enhanced Learned Image Compression
Bowen Hai, Yushun Tang, Zhihai He |
ACCV (7) | 3 |
| 2024 | Dual-Path Adversarial Lifting for Domain Shift Correction in Online Test-Time Adaptation
Yushun Tang, Shuoshuo Chen, Zhihe Lu, Xinchao Wang, Zhihai He |
ECCV (67) | 1 |
| 2024 | Learning Inference-Time Drift Sensor-Actuator for Domain GeneralizationabstractIn machine learning tasks, models trained in the source domain often suffer from performance degradation in the target domain due to domain drift or distribution shift. In this paper, we explore the concept of sensor-actuator design in adaptive control to address this domain drift problem and develop a new approach, called learning inference-time drift sensor-actuator (LIDSA) for domain generalization. The drift sensor network consists of a constraint network and a data converter. The constraint network is learned to extract a set of constraints in the source domain and sense the domain drift by detecting the deviation from these constraints, called constraint error, which is correlated with the classification error. The data converter network then maps this constraint error into an effective guidance signal, which can guide the actuator network to adjust the feature to achieve improved discrimination power and better generalization performance. Our extensive experimental results demonstrate that the proposed LIDSA approach improves the performance of domain generalization over the baseline method. Shuoshuo Chen, Yushun Tang, Zhehan Kan, Zhihai He |
ICASSP | 2 |
| 2024 | Domain-Conditioned Transformer for Fully Test-time Adaptation
Yushun Tang, Shuoshuo Chen, Jiyuan Jia, Yi Zhang 0109, Zhihai He |
ACM Multimedia | 1 |
| 2024 | Training-Free Feature Reconstruction with Sparse Optimization for Vision-Language ModelsabstractIn this paper, we address the challenge of adapting vision-language models (VLMs) to few-shot image recognition in a training-free manner. We observe that existing methods are not able to effectively characterize the semantic relationship between support and query samples in a training-free setting. We recognize that, in the semantic feature space, the feature of the query image is a linear and sparse combination of support image features since support-query pairs are from the class and share the same small set of distinctive visual attributes. Motivated by this interesting observation, we propose a novel method called Training-free Feature ReConstruction with Sparse optimization (TaCo), which formulates the few-shot image recognition task as a feature reconstruction and sparse optimization problem. Specifically, we exploit the VLM to encode the query and support images into features. We utilize sparse optimization to reconstruct the query feature from the corresponding support features. The feature reconstruction error is then used to define the reconstruction similarity. Coupled with the text-image similarity provided by the VLM, our reconstruction similarity analysis accurately characterizes the relationship between support and query images. This results in significantly improved performance in few-shot image recognition. Our extensive experimental results on few-shot recognition demonstrate that our method outperforms existing state-of-the-art approaches by substantial margins. Yi Zhang 0109, Ke Yu 0004, Angelica I. Avilés-Rivero, Jiyuan Jia, Yushun Tang, Zhihai He |
ACM Multimedia | 5 |
| 2024 | Cross-Modal Concept Learning and Inference for Vision-Language Models
Yi Zhang 0109, Ce Zhang 0009, Yushun Tang, Zhihai He |
Neurocomputing | 3 |
| 2023 | BDC-Adapter: Brownian Distance Covariance for Better Vision-Language Reasoning
Yi Zhang 0109, Ce Zhang 0009, Yushun Tang, Zhihai He |
BMVC | 4 |
| 2023 | Self-Correctable and Adaptable Inference for Generalizable Human Pose EstimationabstractA central challenge in human pose estimation, as well as in many other machine learning and prediction tasks, is the generalization problem. The learned network does not have the capability to characterize the prediction error, generate feedback information from the test sample, and correct the prediction error on the fly for each individual test sample, which results in degraded performance in generalization. In this work, we introduce a self-correctable and adaptable inference (SCAI) method to address the generalization challenge of network prediction and use human pose estimation as an example to demonstrate its effectiveness and performance. We learn a correction network to correct the prediction result conditioned by a fitness feedback error. This feedback error is generated by a learned fitness feedback network which maps the prediction result to the original input domain and compares it against the original input. Interestingly, we find that this self-referential feedback error is highly correlated with the actual prediction error. This strong correlation suggests that we can use this error as feedback to guide the correction process. It can be also used as a loss function to quickly adapt and optimize the correction network during the inference process. Our extensive experimental results on human pose estimation demonstrate that the proposed SCAI method is able to significantly improve the generalization capability and performance of human pose estimation. Zhehan Kan, Shuoshuo Chen, Ce Zhang 0009, Yushun Tang, Zhihai He |
CVPR | 4 |
| 2023 | Neuro-Modulated Hebbian Learning for Fully Test-Time AdaptationabstractFully test-time adaptation aims to adapt the network model based on sequential analysis of input samples during the inference stage to address the cross-domain performance degradation problem of deep neural networks. We take inspiration from the biological plausibility learning where the neuron responses are tuned based on a local synapse-change procedure and activated by competitive lateral inhibition rules. Based on these feed-forward learning rules, we design a soft Hebbian learning process which provides an unsupervised and effective mechanism for online adaptation. We observe that the performance of this feed-forward Hebbian learning for fully test-time adaptation can be significantly improved by incorporating a feedback neuromodulation layer. It is able to fine-tune the neuron responses based on the external feedback generated by the error backpropagation from the top inference layers. This leads to our proposed neuro-modulated Hebbian learning (NHL) method for fully test-time adaptation. With the unsupervised feed-forward soft Hebbian learning being combined with a learned neuromodulator to capture feedback from external responses, the source model can be effectively adapted during the testing process. Experimental results on benchmark datasets demonstrate that our proposed method can significantly improve the adaptation performance of network models and outperforms existing state-of-the-art methods. Yushun Tang, Ce Zhang 0009, Shuoshuo Chen, Luziwei Leng, Qinghai Guo, Zhihai He |
CVPR | 1 |
| 2023 | Cross-Inferential Networks for Source-Free Unsupervised Domain AdaptationabstractOne central challenge in source-free unsupervised domain adaptation (UDA) is the lack of an effective approach to evaluate the prediction results of the adapted network model in the target domain. To address this challenge, we propose to explore a new method called cross-inferential networks (CIN). Our main idea is that, when we adapt the network model to predict the sample labels from encoded features, we use these prediction results to construct new training samples with derived labels to learn a new examiner network that performs a different but compatible task in the target domain. Specifically, in this work, the base network model is performing image classification while the examiner network is tasked to perform relative ordering of triplets of samples whose training labels are carefully constructed from the prediction results of the base network model. Two similarity measures, cross-network correlation matrix similarity and attention consistency, are then developed to provide important guidance for the UDA process. Our experimental results on benchmark datasets demonstrate that our proposed CIN approach can significantly improve the performance of source-free UDA. Yushun Tang, Qinghai Guo, Zhihai He |
ICIP | 1 |