EDBT 2026 Demo / reviewers in the wild / expert
Shuai Zhang 0014
dblp:71/208-14
· DBLP profile ↗
20ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-1094-887XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AStar: Boosting Multimodal Reasoning with Automated Structured ThinkingabstractMultimodal large language models excel across diverse domains but struggle with complex visual reasoning tasks. To enhance their reasoning capabilities, current approaches typically rely on explicit search or post-training techniques. However, search-based methods suffer from computational inefficiency due to extensive solution space exploration, while post-training methods demand substantial data, computational resources, and often exhibit training instability. To address these challenges, we propose **AStar**, a training-free, **A**utomatic **S**tructured **t**hinking paradigm for multimod**a**l **r**easoning. Specifically, we introduce novel "thought cards", a lightweight library of high-level reasoning patterns abstracted from prior samples. For each test problem, AStar adaptively retrieves the optimal thought cards and seamlessly integrates these external explicit guidelines with the model’s internal implicit reasoning capabilities. Compared to previous methods, AStar eliminates computationally expensive explicit search and avoids additional complex post-training processes, enabling a more efficient reasoning approach. Extensive experiments demonstrate that our framework achieves 53.9% accuracy on MathVerse (surpassing GPT-4o's 50.2%) and 32.7% on MathVision (outperforming GPT-4o's 30.4%). Further analysis reveals the remarkable transferability of our method: thought cards generated from mathematical reasoning can also be applied to other reasoning tasks, even benefiting general visual perception and understanding. AStar serves as a plug-and-play test-time inference method, compatible with other post-training techniques, providing an important complement to existing multimodal reasoning approaches. Mingkuan Feng, Guocheng Zhai, Shuai Zhang 0014, Zheng Lian 0004, Fangrui Lv, Pengpeng Shao, Ruihan Jin, Zhengqi Wen, Jianhua Tao 0001 |
AAAI | 4 |
| 2026 | ReFL: Reflective Feedback Learning for Hallucination Detection of Large Language ModelsabstractCunhang Fan, Jun Zhang, Xue Zhang, Shuai Zhang, Zhao Lv, Jianhua Tao, Zhengqi Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cunhang Fan, Shuai Zhang 0014, Zhao Lv, Jianhua Tao 0001, Zhengqi Wen |
ACL (1) | 4 |
| 2026 | Two-Stage Regularization-Based Structured Pruning for LLMsabstractMingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mingkuan Feng, Shuai Zhang 0014, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao 0001 |
ACL (1) | 4 |
| 2026 | Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language ModelsabstractJinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che, Zhengqi Wen, Chonghua Liao, Ling Yang, Haoran Luo, Zheng Lian, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mingkuan Feng, Shuai Zhang 0014, Feihu Che, Zhengqi Wen, Chonghua Liao, Zheng Lian 0004, Jianhua Tao 0001 |
ACL (1) | 3 |
| 2026 | SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic LearningabstractReinforcement learning has empowered large language models to act as intelligent agents, yet training them for long-horizon tasks remains challenging due to the scarcity of highquality trajectories, especially under limited resources.Existing methods typically scale up rollout sizes and indiscriminately allocate computational resources among intermediate steps.Such attempts inherently waste substantial computation budget on trivial steps while failing to guarantee sample quality.To address this, we propose SPARK (Strategic Policy-Aware ex-ploRation via Key-state dynamic branching), a novel framework that selectively branches at critical decision states for resource-efficient exploration.Our key insight is to activate adaptive branching exploration at critical decision points to probe promising trajectories, thereby achieving precise resource allocation that prioritizes sampling quality over blind coverage.This design leverages the agent's intrinsic decisionmaking signals to reduce dependence on human priors, enabling the agent to autonomously expand exploration and achieve stronger generalization.Experiments across diverse tasks (e.g., embodied planning), demonstrate that SPARK achieves superior success rates with significantly fewer training samples, exhibiting robust generalization even in unseen scenarios.Our code and checkpoints are available at https://github.com/jinyangwu/SPARK. Shuai Zhang 0014, Zhengqi Wen, Jianhua Tao 0001 |
ACL (1) | 4 |
| 2025 | Code-switching Mediated Sentence-level Semantic LearningabstractCode-switching is a linguistic phenomenon in which different languages are used interactively during conversation. It poses significant performance challenges to natural language processing (NLP) tasks due to the often monolingual nature of the underlying system. We focus on sentence-level semantic associations between the different code-switching expressions. And we propose an innovative task-free semantic learning method based on the semantic property. Specifically, there are many different ways of languages switching for a sentence with the same meaning. We refine this into a semantic computational method by designing the loss of semantic invariant constraint during the model optimization. In this work, we conduct thorough experiments on speech recognition, speech translation, and language modeling tasks. The experimental results fully demonstrate that the proposed method can widely improve the performance of code-switching related tasks. Shuai Zhang 0014, Jiangyan Yi, Zhengqi Wen, Jianhua Tao 0001, Feihu Che, Ruibo Fu |
AAAI | 1 |
| 2025 | Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsabstractRetrieval-Augmented Generation (RAG) has emerged as a key method to address hallucinations in large language models (LLMs).While recent research has extended RAG models to complex noisy scenarios, these explorations often confine themselves to limited noise types and presuppose that noise is inherently detrimental to LLMs, potentially deviating from real-world retrieval environments and restricting practical applicability.In this paper, we define seven distinct noise types from a linguistic perspective and establish a Noise RAG Benchmark (NoiserBench), a comprehensive evaluation framework encompassing multiple datasets and reasoning tasks.Through empirical evaluation of eight representative LLMs with diverse architectures and scales, we reveal that these noises can be further categorized into two practical groups: noise that is beneficial to LLMs (aka beneficial noise) and noise that is harmful to LLMs (aka harmful noise).While harmful noise generally impairs performance, beneficial noise may enhance several aspects of model capabilities and overall performance.Our analysis offers insights for developing robust RAG solutions and mitigating hallucinations across diverse retrieval scenarios.Code is available at Shuai Zhang 0014, Feihu Che, Mingkuan Feng, Pengpeng Shao, Jianhua Tao 0001 |
ACL (1) | 2 |
| 2023 | Detection of Cross-Dataset Fake Audio Based on Prosodic and Pronunciation Features
Chenglong Wang 0001, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Shuai Zhang 0014, Xun Chen 0001 |
INTERSPEECH | 5 |
| 2023 | TO-Rawnet: Improving RawNet with TCN and Orthogonal Regularization for Fake Audio Detection
Chenglong Wang 0001, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Shuai Zhang 0014, Ruibo Fu, Xun Chen 0001 |
INTERSPEECH | 5 |
| 2022 | ADD 2022: the first Audio Deep Synthesis Detection ChallengeabstractAudio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks. Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001 |
ICASSP | 13 |
| 2022 | reducing multilingual context confusion for end-to-end code-switching automatic speech recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Jianhua Tao 0001, Yu Ting Yeung, Liqun Deng |
INTERSPEECH | 1 |
| 2022 | Hybrid Autoregressive and Non-Autoregressive Transformer Models for Speech RecognitionabstractThe autoregressive (AR) models, such as attention-based encoder-decoder models and RNN-Transducer, have achieved great success in speech recognition. They predict the output sequence conditioned on the previous tokens and acoustic encoded states, which is inefficient on GPUs. The non-autoregressive (NAR) models can get rid of the temporal dependency between the output tokens and predict the entire output tokens in one inference step. However, the NAR model still faces two major problems. Firstly, there is still a great gap in performance between the NAR models and the advanced AR models. Secondly, it’s difficult for most of the NAR models to train and converge. We propose a hybrid autoregressive and non-autoregressive transformer (HANAT) model, which integrates AR and NAR models deeply by sharing parameters. We assume that the AR model will assist the NAR model to learn some linguistic dependencies and accelerate the convergence. Furthermore, the two-stage hybrid inference is applied to improve the model performance. All the experiments are conducted on a mandarin dataset ASIEHLL-1 and a english dataset librispeech-960 h. The results show that the HANAT can achieve a competitive performance with the AR model and outperform many complicated NAR models. Besides, the RTF is only 1/5 of the AR model. Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen |
IEEE Signal Process. Lett. | 4 |
| 2021 | Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech RecognitionabstractDespite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models’ performance. In this paper, we propose a decoupled transformer model to use mono-lingual paired data and unpaired text data to alleviate the problem of code-switching data shortage. The model is decoupled into two parts: audio-to-phoneme (A2P) network and phoneme-to-text (P2T) network. The A2P network can learn acoustic pattern scenarios using large-scale monolingual paired data. Meanwhile, it generates multiple phoneme sequence candidates for single audio data in real time during the training process. Then the generated phoneme-text paired data is used to train the P2T network. This network can be pre-trained with large amounts of external unpaired text data. By using monolingual data and unpaired text data, the decoupled transformer model reduces the high dependency on code-switching paired training data of E2E model to a certain extent. Finally, the two networks are optimized jointly through attention fusion. We evaluate the proposed method on the public Mandarin-English code-switching dataset. Compared with our transformer baseline, the proposed method achieves 18.14% relative mix error rate reduction. Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Ye Bai 0001, Jianhua Tao 0001, Zhengqi Wen |
ICASSP | 1 |
| 2021 | End-to-End Spelling Correction Conditioned on Acoustic Feature for Code-Switching Speech Recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Ye Bai 0001, Jianhua Tao 0001, Xuefei Liu, Zhengqi Wen |
Interspeech | 1 |
| 2021 | FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip RegularizationabstractTransducer-based models, such as RNN-Transducer and transformer-transducer, have achieved great success in speech recognition. A typical transducer model decodes the output sequence conditioned on the current acoustic state and previously predicted tokens step by step. Statistically, The number of blank tokens in the prediction results accounts for nearly 90\% of all tokens. It takes a lot of computation and time to predict the blank tokens, but only the non-blank tokens will appear in the final output sequence. Therefore, we propose a method named fast-skip regularization, which tries to align the blank position predicted by a transducer with that predicted by a CTC model. During the inference, the transducer model can predict the blank tokens in advance by a simple CTC project layer without many complicated forward calculations of the transducer decoder and then skip them, which will reduce the computation and improve the inference speed greatly. All experiments are conducted on a public Chinese mandarin dataset AISHELL-1. The results show that the fast-skip regularization can indeed help the transducer model learn the blank position alignments. Besides, the inference with fast-skip can be speeded up nearly 4 times with only a little performance degradation. Zhengkun Tian, Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen |
Interspeech | 5 |
| 2021 | Fast End-to-End Speech Recognition Via Non-Autoregressive Models and Cross-Modal Knowledge Transferring From BERTabstractAttention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because the decoder predicts text tokens (such as characters or words) in an autoregressive manner, it is difficult for an AED model to predict all tokens in parallel. This makes the inference speed relatively slow. In contrast, we propose an end-to-end non-autoregressive speech recognition model called LASO (Listen Attentively, and Spell Once). The model aggregates encoded speech features into the hidden representations corresponding to each token with attention mechanisms. Thus, the model can capture the token relations by self-attention on the aggregated hidden representations from the whole speech signal rather than autoregressive modeling on tokens. Without explicitly autoregressive language modeling, this model predicts all tokens in the sequence in parallel so that the inference is efficient. Moreover, we propose a cross-modal transfer learning method to use a text-modal language model to improve the performance of speech-modal LASO by aligning token semantics. We conduct experiments on two scales of public Chinese speech datasets AISHELL-1 and AISHELL-2. Experimental results show that our proposed model achieves a speedup of about 50× and competitive performance, compared with the autoregressive transformer models. And the cross-modal knowledge transferring from the text-modal model can improve the performance of the speech-modal model. Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen, Shuai Zhang 0014 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | Integrating Knowledge Into End-to-End Speech Recognition From External Text-Only DataabstractAttention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained with speech-text paired data. It is challenging to incorporate external text-only data into AED models. Another issue of the AED model is that it does not use the right context of a text token while predicting the token. To alleviate the above two issues, we propose a unified method called LST (Learn Spelling from Teachers) to integrate knowledge into an AED model from the external text-only data and leverage the whole context in a sentence. The method is divided into two stages. First, in the representation stage, a language model is trained on the text. It can be seen as that the knowledge in the text is compressed into the LM. Then, at the transferring stage, the knowledge is transferred to the AED model via teacher-student learning. To further use the whole context of the text sentence, we propose an LM called causal cloze completer (COR), which estimates the probability of a token, given both the left context and the right context of it. Therefore, with LST training, the AED model can leverage the whole context in the sentence. Different from fusion based methods, which use LM during decoding, the proposed method does not increase any extra complexity at the inference stage. We conduct experiments on two scales of public Chinese datasets AISHELL-1 and AISHELL-2. The experimental results demonstrate the effectiveness of leveraging external text-only data and the whole context in a sentence with our proposed method, compared with baseline hybrid systems and AED model based systems. Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Zhengkun Tian, Shuai Zhang 0014 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Synchronous Transformers for end-to-end Speech RecognitionabstractFor most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition. In this paper, we propose a model named synchronous transformer to address this problem, which can predict the output sequence chunk by chunk. Once a fixed-length chunk of the input sequence is processed by the encoder, the decoder begins to predict symbols immediately. During training, a forward-backward algorithm is introduced to optimize all the possible alignment paths. Our model is evaluated on a Mandarin dataset AISHELL-1. The experiments show that the synchronous transformer is able to perform encoding and decoding synchronously, and achieves a character error rate of 8.91% on the test set. Zhengkun Tian, Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen |
ICASSP | 5 |
| 2020 | Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech RecognitionabstractAlthough attention based end-to-end models have achieved promising performance in speech recognition, the multi-pass forward computation in beam-search increases inference time cost, which limits their practical applications.To address this issue, we propose a non-autoregressive end-to-end speech recognition system called LASO (listen attentively, and spell once).Because of the non-autoregressive property, LASO predicts a textual token in the sequence without the dependence on other tokens.Without beam-search, the one-pass propagation much reduces inference time cost of LASO.And because the model is based on the attention based feedforward structure, the computation can be implemented in parallel efficiently.We conduct experiments on publicly available Chinese dataset AISHELL-1.LASO achieves a character error rate of 6.4%, which outperforms the state-of-the-art autoregressive transformer model (6.7%).The average inference latency is 21 ms, which is 1/50 of the autoregressive transformer model. Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen, Shuai Zhang 0014 |
INTERSPEECH | 6 |
| 2020 | Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech RecognitionabstractNon-autoregressive transformer models have achieved extremely fast inference speed and comparable performance with autoregressive sequence-to-sequence models in neural machine translation.Most of the non-autoregressive transformers decode the target sequence from a predefined-length mask sequence.If the predefined length is too long, it will cause a lot of redundant calculations.If the predefined length is shorter than the length of the target sequence, it will hurt the performance of the model.To address this problem and improve the inference speed, we propose a spike-triggered non-autoregressive transformer model for end-to-end speech recognition, which introduces a CTC module to predict the length of the target sequence and accelerate the convergence.All the experiments are conducted on a public Chinese mandarin dataset AISHELL-1.The results show that the proposed model can accurately predict the length of the target sequence and achieve a competitive performance with the advanced transformers.What's more, the model even achieves a real-time factor of 0.0056, which exceeds all mainstream speech recognition models. Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Shuai Zhang 0014, Zhengqi Wen |
INTERSPEECH | 5 |