Shuanghao Bai

dblp:364/7251 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0002-6047-0242ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 18% Transfer learning and domain adaptation · 15% Representation and self-supervised learning · 11%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning
0.912025
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation · ICML 2025
Machine learning › Representation and self-supervised learning
information bottleneck
0.912025
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation · ICML 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.912025
VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation · ICLR 2025
Machine learning › Reinforcement learning › imitation learning › learning from observation
visual imitation learning
0.912025
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.812024
Jacobian Regularizer-based Neural Granger Causality · ICML 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
0.812024
Soft Prompt Generation for Domain Generalization · ECCV (3) 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
granger causality
0.812024
Jacobian Regularizer-based Neural Granger Causality · ICML 2024
Natural language and speech › Language models and text generation
prompt tuning
0.812024
Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024
Machine learning › Deep learning architectures and training
regularization
0.812024
Jacobian Regularizer-based Neural Granger Causality · ICML 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.812024
Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024
Computer vision › Vision and language
vision-language model
0.812024
Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
domain-invariant representation
0.212024
Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation · AAAI 2024

Methods — techniques the papers use, named apart from their topics

prompt tuning · 1.5speech instruction tuning · 0.9retrieval-augmented generation · 0.9mutual information · 0.9information bottleneck · 0.9vision-language model · 0.8neural network · 0.8jacobian regularizer · 0.8image-guided feature tuning · 0.8feature bank · 0.8
YearPublicationVenuePosition
2025 PromptTA: Prompt-driven Text Adapter for Source-free Domain Generalization
abstract
Source-free domain generalization (SFDG) tackles the challenge of adapting models to unseen target domains without access to source domain data. To deal with this challenging task, recent advances in SFDG have primarily focused on leveraging the text modality of vision-language models such as CLIP. These methods involve developing a transferable linear classifier based on diverse style features extracted from the text and learned prompts or deriving domain-unified text representations from domain banks. However, both style features and domain banks have limitations in capturing comprehensive domain knowledge. In this work, we propose Prompt-Driven Text Adapter (PromptTA) method, which is designed to better capture the distribution of style features and employ resampling to ensure thorough coverage of domain knowledge. To further leverage this rich domain information, we introduce a text adapter that learns from these style features for efficient domain information storage. Extensive experiments conducted on four benchmark datasets demonstrate that PromptTA achieves state-of-the-art performance. The code is available at https://github.com/zhanghr2001/PromptTA.
Shuanghao Bai, Wanqi Zhou, Jingwen Fu, Badong Chen
ICASSP2
2025 VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation
abstract
Vision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human interaction. A typical approach of incorporating speech modality into VLA necessitates a separate speech recognition system to transcribe spoken instructions into text. Such a cascading pipeline raises two major concerns for robotic systems. First, the entire model grows in size and complexity, potentially resulting in redundant computations and increased memory consumption. Second, the transcription procedure would lose non-semantic information in the raw speech, such as voiceprint, which is crucial for a robot to successfully understand and complete customized tasks. To this end, we propose VLAS, the fisrt end-to-end policy model that seamlessly integrates speech modality for robot manipulation. We present a three-stage speech instruction tuning strategy leveraging multimodal datasets, including our manually curated SQA and CSI datasets. Furthermore, to facilitate personalized operations, we develop a voice retrieval-augmented generation (RAG) approach to enhance the robot's performance in tasks requiring individual-specific knowledge. Experimental results show that the proposed VLAS, following either textual or speech instructions, can achieve performance comparable to traditional VLAs on the CALVIN benchmark. In addition, we created a benchmark consisting of customization tasks, where our VLAS demonstrates absolute superiority by fully leveraging the auxiliary information in speech.
Pengxiang Ding, Min Zhang 0068, Zhefei Gong, Shuanghao Bai, Han Zhao 0008
ICLR5
2025 Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
abstract
Behavior Cloning (BC) is a widely adopted visual imitation learning method in robot manipulation. Current BC approaches often enhance generalization by leveraging large datasets and incorporating additional visual and textual modalities to capture more diverse information. However, these methods overlook whether the learned representations contain redundant information and lack a solid theoretical foundation to guide the learning process. To address these limitations, we adopt an information-theoretic perspective and introduce mutual information to quantify and mitigate redundancy in latent representations. Building on this, we incorporate the Information Bottleneck (IB) principle into BC, which extends the idea of reducing redundancy by providing a structured framework for compressing irrelevant information while preserving task-relevant features. This work presents the first comprehensive study on redundancy in latent representations across various methods, backbones, and experimental settings, while extending the generalizability of the IB to BC. Extensive experiments and analyses on the CortexBench and LIBERO benchmarks show consistent performance improvements with IB across various settings, underscoring the importance of reducing input data redundancy and highlighting its practical value for real-world applications.
Shuanghao Bai, Wanqi Zhou, Pengxiang Ding, Badong Chen
ICML1
2025 An information-theoretic approach for heterogeneous differentiable causal discovery
Wanqi Zhou, Shuanghao Bai, Yuqing Xie 0002, Yicong He, Qibin Zhao, Badong Chen
Neural Networks2
2024 Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation
abstract
Recently, despite the unprecedented success of large pre-trained visual-language models (VLMs) on a wide range of downstream tasks, the real-world unsupervised domain adaptation (UDA) problem is still not well explored. Therefore, in this paper, we first experimentally demonstrate that the unsupervised-trained VLMs can significantly reduce the distribution discrepancy between source and target domains, thereby improving the performance of UDA. However, a major challenge for directly deploying such models on downstream UDA tasks is prompt engineering, which requires aligning the domain knowledge of source and target domains, since the performance of UDA is severely influenced by a good domain-invariant representation. We further propose a Prompt-based Distribution Alignment (PDA) method to incorporate the domain knowledge into prompt learning. Specifically, PDA employs a two-branch prompt-tuning paradigm, namely base branch and alignment branch. The base branch focuses on integrating class-related representation into prompts, ensuring discrimination among different classes. To further minimize domain discrepancy, for the alignment branch, we construct feature banks for both the source and target domains and propose image-guided feature tuning (IFT) to make the input attend to feature banks, which effectively integrates self-enhanced and cross-domain features into the model. In this way, these two branches can be mutually promoted to enhance the adaptation of VLMs for UDA. We conduct extensive experiments on three benchmarks to demonstrate that our proposed PDA achieves state-of-the-art performance. The code is available at https://github.com/BaiShuanghao/Prompt-based-Distribution-Alignment.
Shuanghao Bai, Min Zhang 0068, Wanqi Zhou, Siteng Huang, Zhirong Luan, Badong Chen
AAAI1
2024 Soft Prompt Generation for Domain Generalization
Shuanghao Bai, Yuedi Zhang, Wanqi Zhou, Zhirong Luan, Badong Chen
ECCV (3)1
2024 Improving Cross-Domain Few-Shot Classification with Multilayer Perceptron
abstract
Cross-domain few-shot classification (CDFSC) is a challenging and tough task due to the significant distribution discrepancies across different domains. To address this challenge, many approaches aim to learn transferable representations. Multilayer perceptron (MLP) has shown its capability to learn transferable representations in various downstream tasks, such as unsupervised image classification and supervised concept generalization. However, its potential in the few-shot settings has yet to be comprehensively explored. In this study, we investigate the potential of MLP to assist in addressing the challenges of CDFSC. Specifically, we introduce three distinct frameworks incorporating MLP in accordance with three types of few-shot classification methods to verify the effectiveness of MLP. We reveal that MLP can significantly enhance discriminative capabilities and alleviate distribution shifts, which can be supported by our expensive experiments involving 10 baseline models and 12 benchmark datasets. Furthermore, our method even compares favorably against other state-of-the-art CDFSC algorithms.
Shuanghao Bai, Wanqi Zhou, Zhirong Luan, Badong Chen
ICASSP1
2024 Jacobian Regularizer-based Neural Granger Causality
abstract
With the advancement of neural networks, diverse methods for neural Granger causality have emerged, which demonstrate proficiency in handling complex data, and nonlinear relationships. However, the existing framework of neural Granger causality has several limitations. It requires the construction of separate predictive models for each target variable, and the relationship depends on the sparsity on the weights of the first layer, resulting in challenges in effectively modeling complex relationships between variables as well as unsatisfied estimation accuracy of Granger causality. Moreover, most of them cannot grasp full-time Granger causality. To address these drawbacks, we propose a **J**acobian **R**egularizer-based **N**eural **G**ranger **C**ausality (**JRNGC**) approach, a straightforward yet highly effective method for learning multivariate summary Granger causality and full-time Granger causality by constructing a single model for all target variables. Specifically, our method eliminates the sparsity constraints of weights by leveraging an input-output Jacobian matrix regularizer, which can be subsequently represented as the weighted causal matrix in the post-hoc analysis. Extensive experiments show that our proposed approach achieves competitive performance with the state-of-the-art methods for learning summary Granger causality and full-time Granger causality while maintaining lower model complexity and high scalability.
Wanqi Zhou, Shuanghao Bai, Shujian Yu, Qibin Zhao, Badong Chen
ICML2