Zhe Li 0030

dblp:11/751-30 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
23since 2021 · last 2025
0000-0002-0519-7434ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TrInk: Ink Generation with Transformer Network
abstract
Zezhong Jin, Shubhang Desai, Xu Chen, Biyi Fang, Zhuoyi Huang, Zhe Li, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu, Shujie Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Jin, Shubhang Desai, Biyi Fang, Zhuoyi Huang, Zhe Li 0030, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu 0001, Shujie Liu 0001
EMNLP6
2025 Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification
abstract
In recent years, there has been a surge in the use of a pre-trained speech model as a feature extractor for speaker verification (SV). To reduce model complexity, researchers transfer knowledge from a pre-trained model to a lightweight student model, enabling the latter to reach a performance level not attainable by conventional methods. However, due to the differences in model capacity, the student features contain more noise. This results in discrepancies between the teacher and student features at the intermediate layers, negatively impacting feature-level knowledge distillation (KD). To address this issue, we employ a diffusion model to denoise the student features for KD (DenoKD). This approach enables more effective feature-level distillation. Our method, trained with a small ECAPA-TDNN, achieved a 13% improvement over the baseline on the VoxCeleb1-O test set. Further more, the DenoKD mechanism is found to be effective for SV on short test utterances.
Zezhong Jin, Youzhi Tu, Zhe Li 0030, Chong-Xin Gan, Man-Wai Mak
ICASSP3
2025 Spectral-Aware Low-Rank Adaptation for Speaker Verification
abstract
Previous research has shown that the principal singular vectors of a pre-trained model’s weight matrices capture critical knowledge. In contrast, those associated with small singular values may contain noise or less reliable information. As a result, the LoRA-based parameter-efficient fine-tuning (PEFT) approach, which does not constrain the use of the spectral space, may not be effective for tasks that demand high representation capacity. In this study, we enhance existing PEFT techniques by incorporating the spectral information of pre-trained weight matrices into the fine-tuning process. We investigate spectral adaptation strategies with a particular focus on the additive adjustment of top singular vectors. This is accomplished by applying singular value decomposition (SVD) to the pre-trained weight matrices and restricting the fine-tuning within the top spectral space. Extensive speaker verification experiments on VoxCeleb1 and CN-Celeb1 demonstrate enhanced tuning performance with the proposed approach. Code is released at https://github.com/lizhepolyu/SpectralFT.
Zhe Li 0030, Man-Wai Mak, Mert Pilanci, Hung-yi Lee, Helen M. Meng
ICASSP1
2025 Disentangling Speaker and Content in Pre-trained Speech Models with Latent Diffusion for Robust Speaker Verification
Zhe Li 0030, Man-Wai Mak, Jen-Tzung Chien, Mert Pilanci, Zezhong Jin, Helen M. Meng
INTERSPEECH1
2025 IDIR: Identifying and Distilling Informative Relations for Speaker Verification
Chong-Xin Gan, Zhe Li 0030, Zezhong Jin, Man-Wai Mak, Kong-Aik Lee
INTERSPEECH2
2025 Boosting Generalizability in NPC ART Prediction via Multi-omics Feature Mapping
Jiabao Sheng, Zhe Li 0030, Saikit Lam, Jing Cai 0001
MICCAI (15)2
2025 Locate, enhance and fuse: a progressively optimized network for camouflaged object detection
Tianchi Qiu, Zhe Li 0030, Shaokang Ma, Kangwei Liu 0001, Chenyu Zhou 0005
Multim. Tools Appl.3
2024 Dual Parameter-Efficient Fine-Tuning for Speaker Representation Via Speaker Prompt Tuning and Adapters
abstract
Fine-tuning a pre-trained Transformer model (PTM) for speech applications in a parameter-efficient manner offers the dual benefits of reducing memory and leveraging the rich feature representations in massive unlabeled datasets. However, existing parameter-efficient fine-tuning approaches either adapt the classification head or the whole PTM. The former is unsuitable when the PTM is used as a feature extractor, and the latter does not leverage the different degrees of feature abstraction at different Transformer layers. We propose two solutions to address these limitations. First, we apply speaker prompt tuning to update the task-specific embeddings of a PTM. The tuning enhances speaker feature relevance in the speaker embeddings through the cross-attention between prompt and speaker features. Second, we insert adapter blocks into the Transformer encoders and their outputs. This novel arrangement enables the fine-tuned PTM to determine the most suitable layers to extract relevant information for the downstream task. Extensive speaker verification experiments on Voxceleb and CU-MARVEL demonstrate higher parameter efficiency and better model adaptability of the proposed methods than the existing ones.
Zhe Li 0030, Man-Wai Mak, Helen M. Meng
ICASSP1
2024 Dual Guidance Enhancing Camouflaged Object Detection via Focusing Boundary and Localization Representation
abstract
Camouflaged object detection (COD) aims to segment objects that blend into their surrounding environment. However, low-level features in the shallow layers of neural networks, although rich in edge information, often contain a significant amount of redundant information, making it difficult to represent boundary details accurately. On the other hand, deep high-level features retain semantic information for object localization, but the gradual decrease in resolution can introduce biases in representing localization information. To address this issue, we propose a novel boundary and localization representation network (BLR-Net) that guides high-level features to focus on representing localization information while directing low-level features to emphasize boundary details. Firstly, we propose a multi-scale enhanced feature module (MEFM) to capture multi-scale information from backbone features and obtain aggregated feature representations. Next, we propose an extraction boundary module (EBM) that models object boundary features, providing essential boundary information. Subsequently, we introduce a guided learning module (GLM) that utilizes localization features to guide high-level features toward localization representation learning and boundary features to guide low-level features toward boundary representation learning. Finally, we propose a cross-level feature fusion module (CFFM) that aggregates contextual semantic information and gradually fuses multi-level fusion features from the bottom to the top to predict camouflaged objects. Extensive experiments on four benchmark COD datasets demonstrate that BLR-Net outperforms other state-of-the-art COD models.
Zhe Li 0030, Hongbing Ma, Jiabao Sheng
ICME3
2024 Prompt Fusion Interaction Transformer For Aspect-Based Multimodal Sentiment Analysis
abstract
Aspect-based multimodal sentiment analysis (ABMSA) is a recent and popular research area that uses multiple modalities like text and images to determine the sentiment orientation of opinion entities. The main challenge in multimodal sentiment analysis is dynamically modeling each modality and effectively fusing information across different modalities. Existing methods have not considered fine-grained texture features in images, and direct fusion introduces irrelevant and ineffective features unrelated to sentiment information. To address these limitations, we propose a new model for multimodal sentiment analysis, the multimodal prompt fusion interaction Transformer (MPFIT). We designed two key components: 1) the image assist module (IAM), which leverages the self-attention mechanism and statistical pooling to obtain weighted mean and standard deviation vectors, enabling the model to focus on image texture information and reduce image noise. 2) Multimodal prompt fusion (MPF) restricts multimodal fusion to interactions between small prompt tokens that capture vital information from different modalities, allowing the model to focus on features that are more relevant to sentiment information. Experimental results show that our model outperforms baseline models on two publicly available datasets, Twitter-2015 and Twitter-2017. We conducted ablation experiments to evaluate the impact of our key components.
Zhe Li 0030, Chenyu Zhou 0005
ICME3
2024 Multimodal Rumor Detection via Multimodal Prompt Learning
abstract
Pre-trained vision-language (V-L) models exhibit significant generalization capabilities in detecting rumors. However, their reliance on single-modality prompts—either language or vision—limits their flexibility for dynamic adjustments in both representation spaces during rumor detection. To address these limitations, we propose a multimodal rumor detection framework that uses prompt learning in both the vision and language domains to align their representations better. Inspired by recent advances in efficiently tuning large language models, we introduce a set of trainable parameters in the input space, keeping the model backbone frozen. Additionally, we use distinct prompts at various early stages, which helps progressively model the relationships between features, enhancing comprehensive context learning. Extensive experiments with two real-world multimodal datasets demonstrate our framework’s superior ability to distinguish rumors from facts.
Zhe Li 0030, Chenyu Zhou 0005, Jiabao Sheng
IJCNN3
2024 Boundary-Guided Fusion of Multi-Level Features Network for Camouflaged Object Detection
abstract
Camouflaged objects, exhibiting high similarity with their surroundings, pose a substantial challenge for both humans and machines to detect when concealed within the environment. Existing methods for camouflage object detection (COD) struggle in accurately segmenting the overall structure of camouflaged objects. To address this issue, we propose a novel boundary-guided fusion of multi-level features network (BGFM-Net) for COD. In contrast to existing boundary-guided methods, we pay more attention to addressing the significant imbalance in the pixel quantities between boundary and background features, allowing for a more comprehensive representation of boundary features. BGFM-Net primarily consists of a multi-scale aggregation module (MSAM), a boundary-guided feature module (BFM), and a cross-Level fusion module (CLFM). MSAM effectively integrates contextual semantics at different scales, achieving a powerful and efficient feature representation. BFM adeptly combines edge features while constraining interference from background features, guiding the learning of camouflaged object boundary representation. CLFM integrates multi-level features for predicting camouflaged objects while adaptively adjusting channel weights to emphasize important channels and diminish the impact of less relevant channels for the task. Extensive experiments on three benchmark camouflage datasets demonstrate that our BGFM-Net outperforms other state-of-the-art COD models.
Zhe Li 0030, Jiabao Sheng
IJCNN2
2024 Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Guinan Li, Jiajun Deng, Youjun Chen, Mengzhe Geng, Shujie Hu, Zhe Li 0030, Zengrui Jin, Tianzi Wang, Xurong Xie, Helen M. Meng, Xunying Liu
INTERSPEECH6
2024 Parameter-efficient Fine-tuning of Speaker-Aware Dynamic Prompts for Speaker Verification
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Zhe Li 0030, Man-Wai Mak, Hung-yi Lee, Helen M. Meng
INTERSPEECH1
2024 Enhancing Cross-Modal Alignment in Multimodal Sentiment Analysis via Prompt Learning
Zhe Li 0030, Chenyu Zhou 0005
PRCV (5)3
2024 Enhancing Multimodal Rumor Detection with Statistical Image Features and Modal Alignment via Contrastive Learning
Chenyu Zhou 0005, Zhe Li 0030, Jiabao Sheng, Haoyu Wang 0011
PRICAI (3)3
2024 Knowledge-aware image understanding with multi-level visual representation enhancement for visual question answering
Zhe Li 0030, Wushour Slamu, Yanbing Li
Mach. Learn.2
2023 Discriminative Speaker Representation Via Contrastive Learning with Class-Aware Attention in Angular Space
abstract
The challenges in applying contrastive learning to speaker verification (SV) are that the softmax-based contrastive loss lacks discriminative power and that the hard negative pairs can easily influence learning. To overcome the first challenge, we propose a contrastive learning SV framework incorporating an additive angular margin into the supervised contrastive loss in which the margin improves the speaker representation’s discrimination ability. For the second challenge, we introduce a class-aware attention mechanism through which hard negative samples contribute less significantly to the supervised contrastive loss. We also employed gradient-based multi-objective optimization to balance the classification and contrastive loss. Experimental results on CN-Celeb and Voxceleb1 show that this new learning objective can cause the encoder to find an embedding space that exhibits great speaker discrimination across languages.
Zhe Li 0030, Man-Wai Mak, Helen M. Meng
ICASSP1
2023 Multi-view Contrastive Learning with Additive Margin for Adaptive Nasopharyngeal Carcinoma Radiotherapy Prediction
abstract
The accurate prediction of adaptive radiation therapy (ART) for nasopharyngeal carcinoma (NPC) patients before radiation therapy (RT) is crucial for minimizing toxicity and enhancing patient survival rates. Owing to the complexity of the tumor micro-environment, a single high-resolution image offers only limited insight. Furthermore, the traditional softmax-based loss falls short in quantifying a model’s discriminative power. To address these challenges, we introduce a supervised multi-view contrastive learning approach with an additive margin (MMCon). For each patient, we consider four medical images to form multi-view positive pairs, which supply supplementary information and bolster the representation of medical images. We employ supervised contrastive learning to determine the embedding space, ensuring that NPC samples from the same patient or with the same labels stay in close proximity while NPC samples with different labels are distant. To enhance the discriminative ability of the loss function, we incorporate a margin into the contrastive learning process. Experimental results show that this novel learning objective effectively identifies an embedding space with superior discriminative abilities for NPC images.
Jiabao Sheng, Saikit Lam, Zhe Li 0030, Xinzhi Teng, Yuanpeng Zhang 0001, Jing Cai 0001
ICMR3
2023 Emphasizing Boundary-Positioning and Leveraging Multi-scale Feature Fusion for Camouflaged Object Detection
Zhe Li 0030, Chenyu Zhou 0005, Tianchi Qiu
PRCV (12)3
2023 Knowledge Transfer via Leveraging Teacher-Student Network with Visual Attention to Enhance Atmospheric Sand Image Restoration
Zhe Li 0030
PRCV (11)2
2023 Multimodal Rumor Detection by Using Additive Angular Margin with Class-Aware Attention for Hard Samples
Chenyu Zhou 0005, Zhe Li 0030
PRCV (1)3
2023 Query cost estimation in graph databases via emphasizing query dependencies by using a neural reasoning network
abstract
Summary With the increasing complexity of graph queries, query cost estimation has become a key challenge in graph databases. Accurate estimation results are critical for database administrators or database management systems to perform query processing or optimization tasks. An efficient and accurate estimation model can improve the estimation quality and make the produced results credible. Although learning‐based methods have been applied in query cost estimation, most of them are directed at relational queries and cannot be directly used for graph queries. Furthermore, most estimation approaches focus on the correlations between predicates or columns. The dependencies between query schema and query filter conditions and the correlation between query schema are ignored. In this study, we construct a novel deep learning model composed of reasoning and retrieval processes that can accurately capture the potential logical relationships in graph queries. This solves the above problems to some extent. In addition, we propose a query estimation framework that divides the estimation task into query workload generation, training data collection, feature extraction and encoding, and estimation model construction. The results of the experiment on real‐world datasets show that our estimation model can improve the estimation quality and outperforms other compared deep learning models in terms of estimation accuracy.
Zhenzhen He, Tiquan Gu, Zhe Li 0030, Xusheng Du
Concurr. Comput. Pract. Exp.4