VLDB 2026 Research / reviewers in the wild / expert
Nianwen Si
dblp:196/5047
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-4619-4325ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoLoRA: Boosting LLM-based End-to-end Speech Translation with Mixture of Low-rank ExpertsabstractRecently, End-to-End Speech Translation (E2E-ST) methods leveraging large language models (LLMs) have demonstrated strong generalization capabilities and excellent scalability by integrating pre-trained speech encoders with LLMs, where Low-Rank Adaptation (LoRA) is commonly used for parameter-efficient fine-tuning to reduce training costs. However, LoRA's low-rank assumption often fails in multilingual tasks, as the inherent complexity of cross-lingual semantic relationships and syntactic variations exceeds the representational capacity of low-rank matrices. This leads to parameter conflicts across languages, resulting in suboptimal performance. To address this issue, we propose Mixture of Low-Rank Adaptations (MoLoRA), which integrates the Mixture of Experts (MoE) mechanism with LoRA. MoLoRA effectively enhances the model's expressive capacity while maintaining parameter efficiency during training. Specifically, we treat multiple LoRA modules as low-rank experts and introduce a routing mechanism to dynamically activate language-specific experts. Additionally, shared experts are incorporated and consistently activated to model cross-lingual general knowledge. Furthermore, to enhance the robustness and accuracy of speech representations, we propose a Multi-Granularity Representation Fusion module (MGRF). This module mitigates local distortions in frame-level speech representations caused by noise by fusing frame-level and sentence-level features, thereby providing the LLM with more accurate high-level semantic information. We conduct multilingual experiments on the MuST-C and CoVoST-2 datasets. Our method achieves an average BLEU score of 32.2 across eight language pairs on the MuST-C dataset and an average of 36.3 across three language pairs on the CoVoST-2 dataset, establishing a new state-of-the-art (SOTA) performance. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
AAAI | 3 |
| 2026 | RLRoDA: Enhancing Decompilation Performance with Reinforcement Learning and Read-only Data Augmentation
Chengqi Fu, Nianwen Si |
ICIC (22) | 4 |
| 2026 | Gradient-aware knowledge distillation: Tackling gradient insensitivity through teacher guided gradient scaling
Nianwen Si, Hao Zhang 0109, Weiqiang Zhang 0001, Heyu Chang, Dan Qu 0003 |
Neural Networks | 1 |
| 2025 | MPN: Leveraging Multilingual Patch Neuron for Cross-Lingual Model Editing
Nianwen Si, Heyu Chang, Weiqiang Zhang 0001 |
KSEM (2) | 1 |
| 2024 | A lightweight packet forwarding verification in SDN using sketch
Heyu Chang, Nianwen Si |
Comput. Secur. | 3 |
| 2024 | Improving Speech Translation by Understanding the Speech From Latent CodeabstractDue to data scarcity and modal complexity, the semantic representations extracted by the encoder of end-to-end speech translation (E2E-ST) are often flawed, and its decoder will further produce incorrect semantic alignment between the source speech and the target text based on them, which ultimately impairs translation performance. In contrast to previous research, which focused on how to extract better semantic representations, we focus on how to assist the decoder in performing the translation process in the presence of flawed semantic representations. Specifically, we propose a variational speech translation (VST) framework that leverages latent code containing sentence-level semantic information to aid the decoder in accurately aligning the source speech and target text semantically. By leveraging latent code, VST can compensate for flawed frame-level semantic representations from the encoder and aid the decoder in generating accurate translation text. Our experimental results show that VST can be seamlessly integrated with the current state-of-the-art method, achieving substantial performance improvements. Further analysis and visualization demonstrate that the learned latent code indeed contain rich semantic information and can effectively rectify misalignments in decoder. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Decoupled Non-Parametric Knowledge Distillation for end-to-End Speech TranslationabstractExisting techniques often attempt to make knowledge transfer from a powerful machine translation (MT) to speech translation (ST) model with some elaborate techniques, which often requires transcription as extra input during training. However, transcriptions are not always available, and how to improve the ST model performance without transcription, i.e., data efficiency, has rarely been studied in the literature. In this paper, we propose Decoupled Non-parametric Knowledge Distillation (DNKD) from data perspective to improve the data efficiency. Our method follows the knowledge distillation paradigm. However, instead of obtaining the teacher distribution from a sophisticated MT model, we construct it from a non-parametric datastore via k-Nearest-Neighbor (kNN) retrieval, which removes the dependence on transcription and MT model. Then we decouple the classic knowledge distillation loss into target and non-target distillation to enhance the effect of the knowledge among non-target logits, which is the prominent "dark knowledge". Experiments on MuST-C corpus show that, the proposed method can achieve consistent improvement over the strong baseline without requiring any transcription. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003 |
ICASSP | 2 |
| 2023 | Improving Speech Translation by Cross-Modal Multi-Grained Contrastive LearningabstractThe end-to-end speech translation (E2E-ST) model has gradually become a mainstream paradigm due to its low latency and less error propagation. However, it is non-trivial to train such a model well due to the task complexity and data scarcity. The speech-and-text modality differences result in the E2E-ST model performance usually inferior to the corresponding machine translation (MT) model. Based on the above observation, existing methods often use sharing mechanisms to carry outimplicit knowledge transferby imposing various constraints. However, the final model often performs worse on the MT task than the MT model trained alone, which means that the knowledge transfer ability of this method is also limited. To deal with these problems, we propose the FCCL (Fine- andCoarse- GranularityContrastiveLearning) approach for E2E-ST, which makesexplicit knowledge transferthrough cross-modal multi-grained contrastive learning. A key ingredient of our approach is applying contrastive learning at both sentence- and frame-level to give the comprehensive guide for extracting speech representations containing rich semantic information. In addition, we adopt a simple whitening method to alleviate the representation degeneration in the MT model, which adversely affects contrast learning. Experiments on the MuST-C benchmark show that our proposed approach significantly outperforms the state-of-the-art E2E-ST baselines on all eight language pairs. Further analysis indicates that FCCL can free up its capacity from learning grammatical structure information and force more layers to learn semantic information. Hao Zhang 0109, Nianwen Si, Xukui Yang 0001, Dan Qu 0003, Weiqiang Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Fine-grained visual explanations for the convolutional neural network via class discriminative deconvolution
Nianwen Si, Heyu Chang, Dongning Zhao |
Multim. Tools Appl. | 1 |
| 2021 | Spatial-Channel Attention-Based Class Activation Mapping for Interpreting CNN-Based Image Classification ModelsabstractConvolutional neural network (CNN) has been applied widely in various fields. However, it is always hindered by the unexplainable characteristics. Users cannot know why a CNN-based model produces certain recognition results, which is a vulnerability of CNN from the security perspective. To alleviate this problem, in this study, the three existing feature visualization methods of CNN are analyzed in detail firstly, and a unified visualization framework for interpreting the recognition results of CNN is presented. Here, class activation weight (CAW) is considered as the most important factor in the framework. Then, the different types of CAWs are further analyzed, and it is concluded that a linear correlation exists between them. Finally, on this basis, a spatial-channel attention-based class activation mapping (SCA-CAM) method is proposed. This method uses different types of CAWs as attention weights and combines spatial and channel attentions to generate class activation maps, which is capable of using richer features for interpreting the results of CNN. Experiments on four different networks are conducted. The results verify the linear correlation between different CAWs. In addition, compared with the existing methods, the proposed method SCA-CAM can effectively improve the visualization effect of the class activation map with higher flexibility on network structure. Nianwen Si, Dan Qu 0003, Xiangyang Luo 0001, Heyu Chang |
Secur. Commun. Networks | 1 |
| 2018 | Exploring global sentence representation for graph-based dependency parsing using BLSTM-SCNN
Nianwen Si, Hengjun Wang, Yidong Shan |
Pattern Recognit. Lett. | 1 |