EDBT 2026 Demo / reviewers in the wild / expert
Shuanghao Bai
dblp:364/7251
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0002-6047-0242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 18% Transfer learning and domain adaptation · 15% Representation and self-supervised learning · 11% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning |
0.9 | 1 | 2025 | Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation · ICML 2025 |
Machine learning › Representation and self-supervised learning
information bottleneck |
0.9 | 1 | 2025 | Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation · ICML 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation · ICLR 2025 |
Machine learning › Reinforcement learning › imitation learning › learning from observation
visual imitation learning |
0.9 | 1 | 2025 | Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation · ICML 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.8 | 1 | 2024 | Jacobian Regularizer-based Neural Granger Causality · ICML 2024 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.8 | 1 | 2024 | Soft Prompt Generation for Domain Generalization · ECCV (3) 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
granger causality |
0.8 | 1 | 2024 | Jacobian Regularizer-based Neural Granger Causality · ICML 2024 |
Natural language and speech › Language models and text generation
prompt tuning |
0.8 | 1 | 2024 | Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024 |
Machine learning › Deep learning architectures and training
regularization |
0.8 | 1 | 2024 | Jacobian Regularizer-based Neural Granger Causality · ICML 2024 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.8 | 1 | 2024 | Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024 |
Computer vision › Vision and language
vision-language model |
0.8 | 1 | 2024 | Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024 |
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
domain-invariant representation |
0.2 | 1 | 2024 | Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
prompt tuning · 1.5speech instruction tuning · 0.9retrieval-augmented generation · 0.9mutual information · 0.9information bottleneck · 0.9vision-language model · 0.8neural network · 0.8jacobian regularizer · 0.8image-guided feature tuning · 0.8feature bank · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PromptTA: Prompt-driven Text Adapter for Source-free Domain GeneralizationabstractSource-free domain generalization (SFDG) tackles the challenge of adapting models to unseen target domains without access to source domain data. To deal with this challenging task, recent advances in SFDG have primarily focused on leveraging the text modality of vision-language models such as CLIP. These methods involve developing a transferable linear classifier based on diverse style features extracted from the text and learned prompts or deriving domain-unified text representations from domain banks. However, both style features and domain banks have limitations in capturing comprehensive domain knowledge. In this work, we propose Prompt-Driven Text Adapter (PromptTA) method, which is designed to better capture the distribution of style features and employ resampling to ensure thorough coverage of domain knowledge. To further leverage this rich domain information, we introduce a text adapter that learns from these style features for efficient domain information storage. Extensive experiments conducted on four benchmark datasets demonstrate that PromptTA achieves state-of-the-art performance. The code is available at https://github.com/zhanghr2001/PromptTA. Shuanghao Bai, Wanqi Zhou, Jingwen Fu, Badong Chen |
ICASSP | 2 |
| 2025 | VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot ManipulationabstractVision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human interaction. A typical approach of incorporating speech modality into VLA necessitates a separate speech recognition system to transcribe spoken instructions into text. Such a cascading pipeline raises two major concerns for robotic systems. First, the entire model grows in size and complexity, potentially resulting in redundant computations and increased memory consumption. Second, the transcription procedure would lose non-semantic information in the raw speech, such as voiceprint, which is crucial for a robot to successfully understand and complete customized tasks. To this end, we propose VLAS, the fisrt end-to-end policy model that seamlessly integrates speech modality for robot manipulation. We present a three-stage speech instruction tuning strategy leveraging multimodal datasets, including our manually curated SQA and CSI datasets. Furthermore, to facilitate personalized operations, we develop a voice retrieval-augmented generation (RAG) approach to enhance the robot's performance in tasks requiring individual-specific knowledge. Experimental results show that the proposed VLAS, following either textual or speech instructions, can achieve performance comparable to traditional VLAs on the CALVIN benchmark. In addition, we created a benchmark consisting of customization tasks, where our VLAS demonstrates absolute superiority by fully leveraging the auxiliary information in speech. Pengxiang Ding, Min Zhang 0068, Zhefei Gong, Shuanghao Bai, Han Zhao 0008 |
ICLR | 5 |
| 2025 | Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot ManipulationabstractBehavior Cloning (BC) is a widely adopted visual imitation learning method in robot manipulation. Current BC approaches often enhance generalization by leveraging large datasets and incorporating additional visual and textual modalities to capture more diverse information. However, these methods overlook whether the learned representations contain redundant information and lack a solid theoretical foundation to guide the learning process. To address these limitations, we adopt an information-theoretic perspective and introduce mutual information to quantify and mitigate redundancy in latent representations. Building on this, we incorporate the Information Bottleneck (IB) principle into BC, which extends the idea of reducing redundancy by providing a structured framework for compressing irrelevant information while preserving task-relevant features. This work presents the first comprehensive study on redundancy in latent representations across various methods, backbones, and experimental settings, while extending the generalizability of the IB to BC. Extensive experiments and analyses on the CortexBench and LIBERO benchmarks show consistent performance improvements with IB across various settings, underscoring the importance of reducing input data redundancy and highlighting its practical value for real-world applications. Shuanghao Bai, Wanqi Zhou, Pengxiang Ding, Badong Chen |
ICML | 1 |
| 2025 | An information-theoretic approach for heterogeneous differentiable causal discovery
Wanqi Zhou, Shuanghao Bai, Yuqing Xie 0002, Yicong He, Qibin Zhao, Badong Chen |
Neural Networks | 2 |
| 2024 | Prompt-Based Distribution Alignment for Unsupervised Domain AdaptationabstractRecently, despite the unprecedented success of large pre-trained visual-language models (VLMs) on a wide range of downstream tasks, the real-world unsupervised domain adaptation (UDA) problem is still not well explored. Therefore, in this paper, we first experimentally demonstrate that the unsupervised-trained VLMs can significantly reduce the distribution discrepancy between source and target domains, thereby improving the performance of UDA. However, a major challenge for directly deploying such models on downstream UDA tasks is prompt engineering, which requires aligning the domain knowledge of source and target domains, since the performance of UDA is severely influenced by a good domain-invariant representation. We further propose a Prompt-based Distribution Alignment (PDA) method to incorporate the domain knowledge into prompt learning. Specifically, PDA employs a two-branch prompt-tuning paradigm, namely base branch and alignment branch. The base branch focuses on integrating class-related representation into prompts, ensuring discrimination among different classes. To further minimize domain discrepancy, for the alignment branch, we construct feature banks for both the source and target domains and propose image-guided feature tuning (IFT) to make the input attend to feature banks, which effectively integrates self-enhanced and cross-domain features into the model. In this way, these two branches can be mutually promoted to enhance the adaptation of VLMs for UDA. We conduct extensive experiments on three benchmarks to demonstrate that our proposed PDA achieves state-of-the-art performance. The code is available at https://github.com/BaiShuanghao/Prompt-based-Distribution-Alignment. Shuanghao Bai, Min Zhang 0068, Wanqi Zhou, Siteng Huang, Zhirong Luan, Badong Chen |
AAAI | 1 |
| 2024 | Soft Prompt Generation for Domain Generalization
Shuanghao Bai, Yuedi Zhang, Wanqi Zhou, Zhirong Luan, Badong Chen |
ECCV (3) | 1 |
| 2024 | Improving Cross-Domain Few-Shot Classification with Multilayer PerceptronabstractCross-domain few-shot classification (CDFSC) is a challenging and tough task due to the significant distribution discrepancies across different domains. To address this challenge, many approaches aim to learn transferable representations. Multilayer perceptron (MLP) has shown its capability to learn transferable representations in various downstream tasks, such as unsupervised image classification and supervised concept generalization. However, its potential in the few-shot settings has yet to be comprehensively explored. In this study, we investigate the potential of MLP to assist in addressing the challenges of CDFSC. Specifically, we introduce three distinct frameworks incorporating MLP in accordance with three types of few-shot classification methods to verify the effectiveness of MLP. We reveal that MLP can significantly enhance discriminative capabilities and alleviate distribution shifts, which can be supported by our expensive experiments involving 10 baseline models and 12 benchmark datasets. Furthermore, our method even compares favorably against other state-of-the-art CDFSC algorithms. Shuanghao Bai, Wanqi Zhou, Zhirong Luan, Badong Chen |
ICASSP | 1 |
| 2024 | Jacobian Regularizer-based Neural Granger CausalityabstractWith the advancement of neural networks, diverse methods for neural Granger causality have emerged, which demonstrate proficiency in handling complex data, and nonlinear relationships. However, the existing framework of neural Granger causality has several limitations. It requires the construction of separate predictive models for each target variable, and the relationship depends on the sparsity on the weights of the first layer, resulting in challenges in effectively modeling complex relationships between variables as well as unsatisfied estimation accuracy of Granger causality. Moreover, most of them cannot grasp full-time Granger causality. To address these drawbacks, we propose a **J**acobian **R**egularizer-based **N**eural **G**ranger **C**ausality (**JRNGC**) approach, a straightforward yet highly effective method for learning multivariate summary Granger causality and full-time Granger causality by constructing a single model for all target variables. Specifically, our method eliminates the sparsity constraints of weights by leveraging an input-output Jacobian matrix regularizer, which can be subsequently represented as the weighted causal matrix in the post-hoc analysis. Extensive experiments show that our proposed approach achieves competitive performance with the state-of-the-art methods for learning summary Granger causality and full-time Granger causality while maintaining lower model complexity and high scalability. Wanqi Zhou, Shuanghao Bai, Shujian Yu, Qibin Zhao, Badong Chen |
ICML | 2 |