EDBT 2026 Demo / reviewers in the wild / expert
Dan Qu 0003
dblp:54/7732-3
· DBLP profile ↗
32ranked-venue papers
0as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoLoRA: Boosting LLM-based End-to-end Speech Translation with Mixture of Low-rank ExpertsabstractRecently, End-to-End Speech Translation (E2E-ST) methods leveraging large language models (LLMs) have demonstrated strong generalization capabilities and excellent scalability by integrating pre-trained speech encoders with LLMs, where Low-Rank Adaptation (LoRA) is commonly used for parameter-efficient fine-tuning to reduce training costs. However, LoRA's low-rank assumption often fails in multilingual tasks, as the inherent complexity of cross-lingual semantic relationships and syntactic variations exceeds the representational capacity of low-rank matrices. This leads to parameter conflicts across languages, resulting in suboptimal performance. To address this issue, we propose Mixture of Low-Rank Adaptations (MoLoRA), which integrates the Mixture of Experts (MoE) mechanism with LoRA. MoLoRA effectively enhances the model's expressive capacity while maintaining parameter efficiency during training. Specifically, we treat multiple LoRA modules as low-rank experts and introduce a routing mechanism to dynamically activate language-specific experts. Additionally, shared experts are incorporated and consistently activated to model cross-lingual general knowledge. Furthermore, to enhance the robustness and accuracy of speech representations, we propose a Multi-Granularity Representation Fusion module (MGRF). This module mitigates local distortions in frame-level speech representations caused by noise by fusing frame-level and sentence-level features, thereby providing the LLM with more accurate high-level semantic information. We conduct multilingual experiments on the MuST-C and CoVoST-2 datasets. Our method achieves an average BLEU score of 32.2 across eight language pairs on the MuST-C dataset and an average of 36.3 across three language pairs on the CoVoST-2 dataset, establishing a new state-of-the-art (SOTA) performance. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
AAAI | 6 |
| 2026 | Using Knowledge Induction strategies: LLMs can do better in knowledge-driven dialogue tasks
Sisi Peng, Hao Zhang 0109, Shunhang Li, Dan Qu 0003 |
Comput. Speech Lang. | 5 |
| 2026 | QR 3 AG: Dynamic Retrieval-Augmented Generation via Query-Response Relevance Threshold Judgement
Sisi Peng, Shunhang Li, Dan Qu 0003 |
Expert Syst. Appl. | 4 |
| 2026 | MetaAug: Task augmentation via entropy increase for robust meta-learning
Chaolong Hao, Hao Zhang 0109, Dan Qu 0003, Weiqiang Zhang 0001 |
Knowl. Based Syst. | 4 |
| 2026 | Gradient-aware knowledge distillation: Tackling gradient insensitivity through teacher guided gradient scaling
Nianwen Si, Hao Zhang 0109, Weiqiang Zhang 0001, Heyu Chang, Dan Qu 0003 |
Neural Networks | 6 |
| 2026 | SCNNTraffic: A lightweight and energy-efficient traffic classification method based on spiking convolution neural networksabstractEncrypted traffic classification aims to extract effective representations from encrypted network data, whose content remains opaque, to identify applications or user behaviors. Existing methods mainly use pre-trained models for classification, which increases accuracy but requires high computational resources and is difficult to deploy at the edge. This paper introduces an encrypted traffic classification model utilizing Spiking Convolutional Neural Networks (SCNNs) called SCNNTraffic. SCNNTraffic employs SCNNs to capture the time-varying characteristics of network traffic, facilitating smooth and stable feature extraction. This approach achieves effective classification while significantly reducing energy consumption, highlighting its advantages in resource-constrained environments. Moreover, we introduce a cross-gating module based on spiking neurons that facilitates feature fusion and further decreases power usage. Our experimental results demonstrate that the proposed model significantly reduces both the number of network parameters and energy consumption, achieving accuracy of 98.65 $$\%$$ and 92.51 $$\%$$ on the ISCX-VPN and Non-VPN datasets, respectively. Qin Zeng, Dan Qu 0003, Hao Zhang 0109 |
Peer Peer Netw. Appl. | 2 |
| 2026 | Neural Collapse-Based Class-Incremental Learning for Encrypted Traffic Classification
Qin Zeng, Dan Qu 0003, Hao Zhang 0109 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2025 | Think Before Retrieving: Query-Response Relevance Retrieval Augmented GenerationabstractDespite Large Language Models (LLMs) have demonstrated astonishing capabilities across various tasks, they still face limitations when dealing with specialized and knowledge-intensive tasks, such as hallucination, lack of knowledge, or outdated information. Retrieval-Augmented Generation (RAG) can conveniently provide trustworthy, up-to-date external knowledge to large models, effectively mitigating the aforementioned predicaments. However, existing LLMs that employ RAG almost use the retrieved relevant knowledge without careful consideration, overlooking factors like retrieval accuracy and knowledge redundancy, which may interfere with the models and result in lower-quality responses. Therefore, we focus on the study of retrieval necessity. Before retrieving external knowledge, we first assess the relevance of the model’s direct initial output to the input content to determine whether retrieval is needed. Specifically, we construct a retrieval judgment dataset using multi-agent technology and establish threshold criteria for the judgment process based on the Attention-Aware Layer-Wise Relevance Propagation (AttnLRP). We introduce a retrieval judgment stage into the task flow, where a decision on whether to retrieve is made by comparing the threshold criteria with the relevance score of the current round. Experimental results show that, compared to baseline methods, our approach significantly improves both retrieval judgment accuracy and dialogue generation quality. The retrieval judgment accuracy increased by 8.55%, and the responses showed improvements of 4.89 in BLEU and 0.57 in ROUGE scores, effectively enhancing the model’s performance. Sisi Peng, Dan Qu 0003 |
IJCNN | 3 |
| 2025 | Self-consistent Knowledge Generation in Large Language Models: A Unified Framework of Meta-cognitive Prompting and Knowledge Utility Optimization
Sisi Peng, Shunhang Li, Dan Qu 0003 |
PRCV (4) | 4 |
| 2025 | Easy and effective! Data augmentation for knowledge-aware dialogue generation via multi-perspective sentences interaction
Sisi Peng, Dan Qu 0003, Hao Zhang 0109, Shunhang Li, Minchen Xu |
Neurocomputing | 2 |
| 2025 | Behavioral psychology of LLMs: Better task guidance through punishment and reinforcement
Sisi Peng, Shunhang Li, Dan Qu 0003 |
Neurocomputing | 4 |
| 2025 | A Forced Decoding-based Approach for Enhancing Low-resource ASRabstractMultilingual automatic speech recognition represents a crucial research direction in tackling the challenges associated with low-resource scenarios. To effectively incorporate language-specific information in the joint training of a model across multiple languages and leverage linguistic similarities to enhance performance on the target low-resource language, this paper introduces a language similarity evaluation approach based on forced decoding. Specifically, when the target language is specified, the speeches of the source language are decoded into transcription in the target language, and the normalized posterior is utilized as the foundation for evaluating language similarity. Comprehensive experiments and analyses conducted on six low-resource languages reveal that the proposed approach achieves an average word error rate relative reduction of 21.74, 7.68, and 3.45% compared to three widely used benchmark methods, respectively, thereby validating the effectiveness of our approach. Xukui Yang 0001, Dan Qu 0003 |
Neural Process. Lett. | 3 |
| 2025 | TAML-Adapter: Enhancing Adapter Tuning Through Task-Agnostic Meta-Learning for Low-Resource Automatic Speech RecognitionabstractParameter-efficient fine-tuning of pre-trained multilingual speech models can significantly enhance the speech recognition performance of target languages. However, traditional parameter-efficient fine-tuning methods, such as adapter tuning, often face challenges related to random initialization. This can lead to suboptimal performance when adapting to languages with limited resources. To address this issue, this paper introduces TAML-Adapter, which utilizes the Task-Agnostic Meta-Learning algorithm to initialize the parameters of the adapters before fine-tuning in target low-resource languages. Comprehensive experiments conducted on the Common Voice and Fleurs datasets highlight the superior performance of TAML-Adapter in five languages with limited resources. In addition, the TAML-Adapter demonstrates superior generalizability and extensibility compared to similar competing methods. Xukui Yang 0001, Yangli Xi, Dan Qu 0003 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Meta-Adapter for Self-Supervised Speech Models: A Solution to Low-Resource Speech Recognition ChallengesabstractSelf-supervised models have demonstrated remarkable performance in speech processing by learning latent representations from large amounts of unlabeled data. Although these models yield promising results on low-resource languages, the computational expense of fine-tuning all model parameters is prohibitively high. Adapters offer a solution by incorporating lightweight bottleneck structures into pre-trained models, enabling efficient parameter adaptation for downstream tasks. However, randomly initialized adapters often underperform in low-resource scenarios, limiting their applicability in low-resource languages. To address this issue, we develop the Meta-Adapter for self-supervised models to obtain meta-initialized parameters that facilitate quick adaptation to low-resource languages. Extensive experiments on the Common Voice and FLEURS datasets demonstrate the superior performance of Meta-Adapters on 12 low-resource languages spanning four different language families. Moreover, Meta-adapters show better generalization and extensibility than traditional pretraining methods. Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
LREC/COLING | 5 |
| 2024 | Exploring the Potential of Prompting Methods in Low-Resource Speech Recognition with Whisper
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
NLPCC (3) | 5 |
| 2024 | Meta adversarial learning improves low-resource speech recognition
Xukui Yang 0001, Hao Zhang 0109, Dan Qu 0003 |
Comput. Speech Lang. | 5 |
| 2024 | Improving cross-lingual low-resource speech recognition by Task-based Meta PolyLoss
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
Comput. Speech Lang. | 5 |
| 2024 | Meta-Adaptable-Adapter: Efficient adaptation of self-supervised models for low-resource speech recognition
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
Neurocomputing | 5 |
| 2024 | A Lightweight Task-Agreement Meta Learning for Low-Resource Speech RecognitionabstractAbstract Meta-learning has proven to be a powerful paradigm for transferring knowledge from prior tasks to facilitate the quick learning of new tasks in automatic speech recognition. However, the differences between languages (tasks) lead to variations in task learning directions, causing the harmful competition for model’s limited resources. To address this challenge, we introduce the task-agreement multilingual meta-learning (TAMML), which adopts the gradient agreement algorithm to guide the model parameters towards a direction where tasks exhibit greater consistency. However, the computation and storage cost of TAMML grows dramatically with model’s depth increases. To address this, we further propose a simplification called TAMML-Light which only uses the output layer for gradient calculation. Experiments on three datasets demonstrate that TAMML and TAMML-Light achieve outperform meta-learning approaches, yielding superior results.Furthermore, TAMML-Light can reduce at least 80 $$\%$$ % of the relative increased computation expenses compared to TAMML. Hao Zhang 0109, Dan Qu 0003, Xukui Yang 0001 |
Neural Process. Lett. | 4 |
| 2024 | Meta-Prompt: Boosting Whisper's Performance in Low-Resource Speech RecognitionabstractRecent advancements in large-scale pre-trained automatic speech recognition (ASR) foundation models (e.g., Whisper) have exhibited remarkable performance in speech processing tasks. A recently emerging paradigm, prompt tuning, offers a parameter-efficient approach for fine-tuning, which has proven to be effective in enhancing the adaptation of pre-trained models to downstream tasks. In this paper, we first explore the prompting method for low-resource speech recognition based on Whisper. Although effective, it poses a challenge in the few-shot scenario due to its high sensitivity to initialization. To address this problem, we propose a novel meta-prompt for low-resource speech recognition that leverages the benefits of meta-learning for fast learning. Moreover, we further present a lightweight version of meta-prompt that omits the learning of encoder-prompt, reducing computational and storage costs. Extensive experiments on FLEURS datasets demonstrate consistent improvements across eleven target languages, showing better generalizability. Notably, meta-prompt achieves similar performance with a 20%-shot compared to prompt tuning with a 50%-shot setting, suggesting excellent few-shot learning ability. Hao Zhang 0109, Dan Qu 0003 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Improving Speech Translation by Understanding the Speech From Latent CodeabstractDue to data scarcity and modal complexity, the semantic representations extracted by the encoder of end-to-end speech translation (E2E-ST) are often flawed, and its decoder will further produce incorrect semantic alignment between the source speech and the target text based on them, which ultimately impairs translation performance. In contrast to previous research, which focused on how to extract better semantic representations, we focus on how to assist the decoder in performing the translation process in the presence of flawed semantic representations. Specifically, we propose a variational speech translation (VST) framework that leverages latent code containing sentence-level semantic information to aid the decoder in accurately aligning the source speech and target text semantically. By leveraging latent code, VST can compensate for flawed frame-level semantic representations from the encoder and aid the decoder in generating accurate translation text. Our experimental results show that VST can be seamlessly integrated with the current state-of-the-art method, achieving substantial performance improvements. Further analysis and visualization demonstrate that the learned latent code indeed contain rich semantic information and can effectively rectify misalignments in decoder. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Decoupled Non-Parametric Knowledge Distillation for end-to-End Speech TranslationabstractExisting techniques often attempt to make knowledge transfer from a powerful machine translation (MT) to speech translation (ST) model with some elaborate techniques, which often requires transcription as extra input during training. However, transcriptions are not always available, and how to improve the ST model performance without transcription, i.e., data efficiency, has rarely been studied in the literature. In this paper, we propose Decoupled Non-parametric Knowledge Distillation (DNKD) from data perspective to improve the data efficiency. Our method follows the knowledge distillation paradigm. However, instead of obtaining the teacher distribution from a sophisticated MT model, we construct it from a non-parametric datastore via k-Nearest-Neighbor (kNN) retrieval, which removes the dependence on transcription and MT model. Then we decouple the classic knowledge distillation loss into target and non-target distillation to enhance the effect of the knowledge among non-target logits, which is the prominent "dark knowledge". Experiments on MuST-C corpus show that, the proposed method can achieve consistent improvement over the strong baseline without requiring any transcription. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
ICASSP | 6 |
| 2023 | Task-Consistent Meta Learning for Low-Resource Speech Recognition
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
NLPCC (1) | 5 |
| 2023 | Task-based Meta Focal Loss for Multilingual Low-resource Speech RecognitionabstractLow-resource automatic speech recognition is a challenging task due to a lack of labeled training data. To resolve this issue, multilingual meta-learning learns a better model initialization from many source language tasks for fast adaptation to unseen target languages. However, for diverse source languages, the quantity and difficulty vary greatly because of their different data scales and phonological systems. These differences lead to task-quantity and task-difficulty imbalance issues and thus a failure of multilingual meta-learning ASR. In this work, we propose a task-based meta focal loss (TMFL) approach to address this tough challenge. Specifically, we introduce a hard-task moderator and update the meta-parameters using gradients from both the support set and query set. Our proposed approach focuses more on hard tasks and makes full use of the data from hard tasks. Moreover, we analyze the significance of the hard task moderator and interpret its significance at the sample level. Experiment results show that the proposed method, TMFL, significantly outperforms the state-of-the-art multilingual meta-learning on all target languages for the IARPA BABEL and OpenSLR datasets, especially under very-low-resource conditions. In particular, it can reduce character error rate from 72% to 60% by fine-tuning the pre-trained model with about 22 hours of Vietnamese data. Hao Zhang 0109, Dan Qu 0003, Xukui Yang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Improving Speech Translation by Cross-Modal Multi-Grained Contrastive LearningabstractThe end-to-end speech translation (E2E-ST) model has gradually become a mainstream paradigm due to its low latency and less error propagation. However, it is non-trivial to train such a model well due to the task complexity and data scarcity. The speech-and-text modality differences result in the E2E-ST model performance usually inferior to the corresponding machine translation (MT) model. Based on the above observation, existing methods often use sharing mechanisms to carry outimplicit knowledge transferby imposing various constraints. However, the final model often performs worse on the MT task than the MT model trained alone, which means that the knowledge transfer ability of this method is also limited. To deal with these problems, we propose the FCCL (Fine- andCoarse- GranularityContrastiveLearning) approach for E2E-ST, which makesexplicit knowledge transferthrough cross-modal multi-grained contrastive learning. A key ingredient of our approach is applying contrastive learning at both sentence- and frame-level to give the comprehensive guide for extracting speech representations containing rich semantic information. In addition, we adopt a simple whitening method to alleviate the representation degeneration in the MT model, which adversely affects contrast learning. Experiments on the MuST-C benchmark show that our proposed approach significantly outperforms the state-of-the-art E2E-ST baselines on all eight language pairs. Further analysis indicates that FCCL can free up its capacity from learning grammatical structure information and force more layers to learn semantic information. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003, Weiqiang Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | DropDim: A Regularization Method for Transformer NetworksabstractWe introduce DropDim, a structured dropout method designed for regularizing the self-attention mechanism, which is a key component of the transformer. In contrast to the general dropout method, which randomly drops neurons, DropDim drops part of the embedding dimensions. In this way, the semantic information can be completely discarded. Thus, the excessive co-adapting between different embedding dimensions can be broken, and the self-attention is forced to encode meaningful features with a certain number of embedding dimensions erased. Experiments on a wide range of tasks executed on the MUST-C English-Germany dataset show that DropDim can effectively improve model performance, reduce over-fitting, and show complementary effects with other regularization methods. When combined with label smoothing, the WER can be reduced from 19.1% to 15.1% on the ASR task, and the BLEU value can be increased from 26.90 to 28.38 on the MT task. On the ST task, the model can reach a BLEU score of 22.99, an increase by 1.86 BLEU points compared to the strong baseline. Hao Zhang 0109, Dan Qu 0003, Keji Shao, Xukui Yang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Spatial-Channel Attention-Based Class Activation Mapping for Interpreting CNN-Based Image Classification ModelsabstractConvolutional neural network (CNN) has been applied widely in various fields. However, it is always hindered by the unexplainable characteristics. Users cannot know why a CNN-based model produces certain recognition results, which is a vulnerability of CNN from the security perspective. To alleviate this problem, in this study, the three existing feature visualization methods of CNN are analyzed in detail firstly, and a unified visualization framework for interpreting the recognition results of CNN is presented. Here, class activation weight (CAW) is considered as the most important factor in the framework. Then, the different types of CAWs are further analyzed, and it is concluded that a linear correlation exists between them. Finally, on this basis, a spatial-channel attention-based class activation mapping (SCA-CAM) method is proposed. This method uses different types of CAWs as attention weights and combines spatial and channel attentions to generate class activation maps, which is capable of using richer features for interpreting the results of CNN. Experiments on four different networks are conducted. The results verify the linear correlation between different CAWs. In addition, compared with the existing methods, the proposed method SCA-CAM can effectively improve the visualization effect of the class activation map with higher flexibility on network structure. Nianwen Si, Dan Qu 0003, Xiangyang Luo 0001, Heyu Chang |
Secur. Commun. Networks | 3 |
| 2018 | Semi-supervised minimum redundancy maximum relevance feature selection for audio classification
Xukui Yang 0001, Liang He 0003, Dan Qu 0003, Weiqiang Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2016 | The NDSC transcription system for the 2016 multi-genre broadcast challengeabstractThe National Digital Switching System Engineering and Technological R&D Center (NDSC) speech-to-text transcription system for the 2016 multi-genre broadcast challenge is described. Various acoustic models based on deep neural network (DNN), such as hybrid DNN, long short term memory recurrent neural network (LSTM RNN), and time delay neural network (TDNN), are trained. The system also makes use of recurrent neural network language models (RNNLMs) for re-scoring and minimum Bayes risk (MBR) combination. The WER on test dataset of the speech-to-text task is 18.2%. Furthermore, to simulate real applications where manual segmentations were not available an automatic segmentation system based on long-term information is proposed. WERs based on the automatically generated segments were slightly worse than that based on the manual segmentations. Xukui Yang 0001, Dan Qu 0003, Weiqiang Zhang 0001 |
SLT | 2 |
| 2014 | Speaker adaptation based on sparse and low-rank eigenphone matrix estimation
Dan Qu 0003, Weiqiang Zhang 0001, Bi-Cheng Li |
INTERSPEECH | 2 |
| 2013 | Rapid speaker adaptation using compressive sensing
Dan Qu 0003, Weiqiang Zhang 0001, Bi-Cheng Li |
Speech Commun. | 2 |
| 2012 | Bayesian Speaker Adaptation Based on a New Hierarchical Probabilistic ModelabstractIn this paper, a new hierarchical Bayesian speaker adaptation method called HMAP is proposed that combines the advantages of three conventional algorithms, maximum a posteriori (MAP), maximum-likelihood linear regression (MLLR), and eigenvoice, resulting in excellent performance across a wide range of adaptation conditions. The new method efficiently utilizes intra-speaker and inter-speaker correlation information through modeling phone and speaker subspaces in a consistent hierarchical Bayesian way. The phone variations for a specific speaker are assumed to be located in a low-dimensional subspace. The phone coordinate, which is shared among different speakers, implicitly contains the intra-speaker correlation information. For a specific speaker, the phone variation, represented by speaker-dependent eigenphones, are concatenated into a supervector. The eigenphone supervector space is also a low dimensional speaker subspace, which contains inter-speaker correlation information. Using principal component analysis (PCA), a new hierarchical probabilistic model for the generation of the speech observations is obtained. Speaker adaptation based on the new hierarchical model is derived using the maximum a posteriori criterion in a top-down manner. Both batch adaptation and online adaptation schemes are proposed. With tuned parameters, the new method can handle varying amounts of adaptation data automatically and efficiently. Experimental results on a Mandarin Chinese continuous speech recognition task show good performance under all testing conditions. Weiqiang Zhang 0001, Bi-Cheng Li, Dan Qu 0003, Michael T. Johnson |
IEEE Trans. Speech Audio Process. | 4 |