EDBT 2026 Demo / reviewers in the wild / expert
Hao Zhang 0109
dblp:55/2270-109
· DBLP profile ↗
21ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-8852-6311ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoLoRA: Boosting LLM-based End-to-end Speech Translation with Mixture of Low-rank ExpertsabstractRecently, End-to-End Speech Translation (E2E-ST) methods leveraging large language models (LLMs) have demonstrated strong generalization capabilities and excellent scalability by integrating pre-trained speech encoders with LLMs, where Low-Rank Adaptation (LoRA) is commonly used for parameter-efficient fine-tuning to reduce training costs. However, LoRA's low-rank assumption often fails in multilingual tasks, as the inherent complexity of cross-lingual semantic relationships and syntactic variations exceeds the representational capacity of low-rank matrices. This leads to parameter conflicts across languages, resulting in suboptimal performance. To address this issue, we propose Mixture of Low-Rank Adaptations (MoLoRA), which integrates the Mixture of Experts (MoE) mechanism with LoRA. MoLoRA effectively enhances the model's expressive capacity while maintaining parameter efficiency during training. Specifically, we treat multiple LoRA modules as low-rank experts and introduce a routing mechanism to dynamically activate language-specific experts. Additionally, shared experts are incorporated and consistently activated to model cross-lingual general knowledge. Furthermore, to enhance the robustness and accuracy of speech representations, we propose a Multi-Granularity Representation Fusion module (MGRF). This module mitigates local distortions in frame-level speech representations caused by noise by fusing frame-level and sentence-level features, thereby providing the LLM with more accurate high-level semantic information. We conduct multilingual experiments on the MuST-C and CoVoST-2 datasets. Our method achieves an average BLEU score of 32.2 across eight language pairs on the MuST-C dataset and an average of 36.3 across three language pairs on the CoVoST-2 dataset, establishing a new state-of-the-art (SOTA) performance. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
AAAI | 1 |
| 2026 | Using Knowledge Induction strategies: LLMs can do better in knowledge-driven dialogue tasks
Sisi Peng, Hao Zhang 0109, Shunhang Li, Dan Qu 0003 |
Comput. Speech Lang. | 3 |
| 2026 | MetaAug: Task augmentation via entropy increase for robust meta-learning
Chaolong Hao, Hao Zhang 0109, Dan Qu 0003, Weiqiang Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2026 | Gradient-aware knowledge distillation: Tackling gradient insensitivity through teacher guided gradient scaling
Nianwen Si, Hao Zhang 0109, Weiqiang Zhang 0001, Heyu Chang, Dan Qu 0003 |
Neural Networks | 2 |
| 2026 | SCNNTraffic: A lightweight and energy-efficient traffic classification method based on spiking convolution neural networksabstractEncrypted traffic classification aims to extract effective representations from encrypted network data, whose content remains opaque, to identify applications or user behaviors. Existing methods mainly use pre-trained models for classification, which increases accuracy but requires high computational resources and is difficult to deploy at the edge. This paper introduces an encrypted traffic classification model utilizing Spiking Convolutional Neural Networks (SCNNs) called SCNNTraffic. SCNNTraffic employs SCNNs to capture the time-varying characteristics of network traffic, facilitating smooth and stable feature extraction. This approach achieves effective classification while significantly reducing energy consumption, highlighting its advantages in resource-constrained environments. Moreover, we introduce a cross-gating module based on spiking neurons that facilitates feature fusion and further decreases power usage. Our experimental results demonstrate that the proposed model significantly reduces both the number of network parameters and energy consumption, achieving accuracy of 98.65 $$\%$$ and 92.51 $$\%$$ on the ISCX-VPN and Non-VPN datasets, respectively. Qin Zeng, Dan Qu 0003, Hao Zhang 0109 |
Peer Peer Netw. Appl. | 3 |
| 2026 | Neural Collapse-Based Class-Incremental Learning for Encrypted Traffic Classification
Qin Zeng, Dan Qu 0003, Hao Zhang 0109 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | Easy and effective! Data augmentation for knowledge-aware dialogue generation via multi-perspective sentences interaction
Sisi Peng, Dan Qu 0003, Hao Zhang 0109, Shunhang Li, Minchen Xu |
Neurocomputing | 4 |
| 2024 | Meta-Adapter for Self-Supervised Speech Models: A Solution to Low-Resource Speech Recognition ChallengesabstractSelf-supervised models have demonstrated remarkable performance in speech processing by learning latent representations from large amounts of unlabeled data. Although these models yield promising results on low-resource languages, the computational expense of fine-tuning all model parameters is prohibitively high. Adapters offer a solution by incorporating lightweight bottleneck structures into pre-trained models, enabling efficient parameter adaptation for downstream tasks. However, randomly initialized adapters often underperform in low-resource scenarios, limiting their applicability in low-resource languages. To address this issue, we develop the Meta-Adapter for self-supervised models to obtain meta-initialized parameters that facilitate quick adaptation to low-resource languages. Extensive experiments on the Common Voice and FLEURS datasets demonstrate the superior performance of Meta-Adapters on 12 low-resource languages spanning four different language families. Moreover, Meta-adapters show better generalization and extensibility than traditional pretraining methods. Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
LREC/COLING | 2 |
| 2024 | Exploring the Potential of Prompting Methods in Low-Resource Speech Recognition with Whisper
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
NLPCC (3) | 3 |
| 2024 | Meta adversarial learning improves low-resource speech recognition
Xukui Yang 0001, Hao Zhang 0109, Dan Qu 0003 |
Comput. Speech Lang. | 3 |
| 2024 | Improving cross-lingual low-resource speech recognition by Task-based Meta PolyLoss
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
Comput. Speech Lang. | 2 |
| 2024 | Dual Knowledge Distillation for neural machine translation
Yuxian Wan, Hao Zhang 0109, Yanxia Li |
Comput. Speech Lang. | 4 |
| 2024 | Meta-Adaptable-Adapter: Efficient adaptation of self-supervised models for low-resource speech recognition
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
Neurocomputing | 2 |
| 2024 | A Lightweight Task-Agreement Meta Learning for Low-Resource Speech RecognitionabstractAbstract Meta-learning has proven to be a powerful paradigm for transferring knowledge from prior tasks to facilitate the quick learning of new tasks in automatic speech recognition. However, the differences between languages (tasks) lead to variations in task learning directions, causing the harmful competition for model’s limited resources. To address this challenge, we introduce the task-agreement multilingual meta-learning (TAMML), which adopts the gradient agreement algorithm to guide the model parameters towards a direction where tasks exhibit greater consistency. However, the computation and storage cost of TAMML grows dramatically with model’s depth increases. To address this, we further propose a simplification called TAMML-Light which only uses the output layer for gradient calculation. Experiments on three datasets demonstrate that TAMML and TAMML-Light achieve outperform meta-learning approaches, yielding superior results.Furthermore, TAMML-Light can reduce at least 80 $$\%$$ % of the relative increased computation expenses compared to TAMML. Hao Zhang 0109, Dan Qu 0003, Xukui Yang 0001 |
Neural Process. Lett. | 2 |
| 2024 | Meta-Prompt: Boosting Whisper's Performance in Low-Resource Speech RecognitionabstractRecent advancements in large-scale pre-trained automatic speech recognition (ASR) foundation models (e.g., Whisper) have exhibited remarkable performance in speech processing tasks. A recently emerging paradigm, prompt tuning, offers a parameter-efficient approach for fine-tuning, which has proven to be effective in enhancing the adaptation of pre-trained models to downstream tasks. In this paper, we first explore the prompting method for low-resource speech recognition based on Whisper. Although effective, it poses a challenge in the few-shot scenario due to its high sensitivity to initialization. To address this problem, we propose a novel meta-prompt for low-resource speech recognition that leverages the benefits of meta-learning for fast learning. Moreover, we further present a lightweight version of meta-prompt that omits the learning of encoder-prompt, reducing computational and storage costs. Extensive experiments on FLEURS datasets demonstrate consistent improvements across eleven target languages, showing better generalizability. Notably, meta-prompt achieves similar performance with a 20%-shot compared to prompt tuning with a 50%-shot setting, suggesting excellent few-shot learning ability. Hao Zhang 0109, Dan Qu 0003 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Improving Speech Translation by Understanding the Speech From Latent CodeabstractDue to data scarcity and modal complexity, the semantic representations extracted by the encoder of end-to-end speech translation (E2E-ST) are often flawed, and its decoder will further produce incorrect semantic alignment between the source speech and the target text based on them, which ultimately impairs translation performance. In contrast to previous research, which focused on how to extract better semantic representations, we focus on how to assist the decoder in performing the translation process in the presence of flawed semantic representations. Specifically, we propose a variational speech translation (VST) framework that leverages latent code containing sentence-level semantic information to aid the decoder in accurately aligning the source speech and target text semantically. By leveraging latent code, VST can compensate for flawed frame-level semantic representations from the encoder and aid the decoder in generating accurate translation text. Our experimental results show that VST can be seamlessly integrated with the current state-of-the-art method, achieving substantial performance improvements. Further analysis and visualization demonstrate that the learned latent code indeed contain rich semantic information and can effectively rectify misalignments in decoder. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
IEEE Signal Process. Lett. | 1 |
| 2023 | Decoupled Non-Parametric Knowledge Distillation for end-to-End Speech TranslationabstractExisting techniques often attempt to make knowledge transfer from a powerful machine translation (MT) to speech translation (ST) model with some elaborate techniques, which often requires transcription as extra input during training. However, transcriptions are not always available, and how to improve the ST model performance without transcription, i.e., data efficiency, has rarely been studied in the literature. In this paper, we propose Decoupled Non-parametric Knowledge Distillation (DNKD) from data perspective to improve the data efficiency. Our method follows the knowledge distillation paradigm. However, instead of obtaining the teacher distribution from a sophisticated MT model, we construct it from a non-parametric datastore via k-Nearest-Neighbor (kNN) retrieval, which removes the dependence on transcription and MT model. Then we decouple the classic knowledge distillation loss into target and non-target distillation to enhance the effect of the knowledge among non-target logits, which is the prominent "dark knowledge". Experiments on MuST-C corpus show that, the proposed method can achieve consistent improvement over the strong baseline without requiring any transcription. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
ICASSP | 1 |
| 2023 | Task-Consistent Meta Learning for Low-Resource Speech Recognition
Hao Zhang 0109, Xukui Yang 0001, Dan Qu 0003 |
NLPCC (1) | 2 |
| 2023 | Task-based Meta Focal Loss for Multilingual Low-resource Speech RecognitionabstractLow-resource automatic speech recognition is a challenging task due to a lack of labeled training data. To resolve this issue, multilingual meta-learning learns a better model initialization from many source language tasks for fast adaptation to unseen target languages. However, for diverse source languages, the quantity and difficulty vary greatly because of their different data scales and phonological systems. These differences lead to task-quantity and task-difficulty imbalance issues and thus a failure of multilingual meta-learning ASR. In this work, we propose a task-based meta focal loss (TMFL) approach to address this tough challenge. Specifically, we introduce a hard-task moderator and update the meta-parameters using gradients from both the support set and query set. Our proposed approach focuses more on hard tasks and makes full use of the data from hard tasks. Moreover, we analyze the significance of the hard task moderator and interpret its significance at the sample level. Experiment results show that the proposed method, TMFL, significantly outperforms the state-of-the-art multilingual meta-learning on all target languages for the IARPA BABEL and OpenSLR datasets, especially under very-low-resource conditions. In particular, it can reduce character error rate from 72% to 60% by fine-tuning the pre-trained model with about 22 hours of Vietnamese data. Hao Zhang 0109, Dan Qu 0003, Xukui Yang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Improving Speech Translation by Cross-Modal Multi-Grained Contrastive LearningabstractThe end-to-end speech translation (E2E-ST) model has gradually become a mainstream paradigm due to its low latency and less error propagation. However, it is non-trivial to train such a model well due to the task complexity and data scarcity. The speech-and-text modality differences result in the E2E-ST model performance usually inferior to the corresponding machine translation (MT) model. Based on the above observation, existing methods often use sharing mechanisms to carry outimplicit knowledge transferby imposing various constraints. However, the final model often performs worse on the MT task than the MT model trained alone, which means that the knowledge transfer ability of this method is also limited. To deal with these problems, we propose the FCCL (Fine- andCoarse- GranularityContrastiveLearning) approach for E2E-ST, which makesexplicit knowledge transferthrough cross-modal multi-grained contrastive learning. A key ingredient of our approach is applying contrastive learning at both sentence- and frame-level to give the comprehensive guide for extracting speech representations containing rich semantic information. In addition, we adopt a simple whitening method to alleviate the representation degeneration in the MT model, which adversely affects contrast learning. Experiments on the MuST-C benchmark show that our proposed approach significantly outperforms the state-of-the-art E2E-ST baselines on all eight language pairs. Further analysis indicates that FCCL can free up its capacity from learning grammatical structure information and force more layers to learn semantic information. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003, Weiqiang Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | DropDim: A Regularization Method for Transformer NetworksabstractWe introduce DropDim, a structured dropout method designed for regularizing the self-attention mechanism, which is a key component of the transformer. In contrast to the general dropout method, which randomly drops neurons, DropDim drops part of the embedding dimensions. In this way, the semantic information can be completely discarded. Thus, the excessive co-adapting between different embedding dimensions can be broken, and the self-attention is forced to encode meaningful features with a certain number of embedding dimensions erased. Experiments on a wide range of tasks executed on the MUST-C English-Germany dataset show that DropDim can effectively improve model performance, reduce over-fitting, and show complementary effects with other regularization methods. When combined with label smoothing, the WER can be reduced from 19.1% to 15.1% on the ASR task, and the BLEU value can be increased from 26.90 to 28.38 on the MT task. On the ST task, the model can reach a BLEU score of 22.99, an increase by 1.86 BLEU points compared to the strong baseline. Hao Zhang 0109, Dan Qu 0003, Keji Shao, Xukui Yang 0001 |
IEEE Signal Process. Lett. | 1 |