Yuke Li 0002

dblp:128/0927-2 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
15since 2021 · last 2025
0009-0003-0935-2483ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 15 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
abstract
Recently, the Mixture of Expert (MoE) architecture, such as LR-MoE, is often used to alleviate the impact of language confusion on the multilingual ASR (MASR) task. However, it still faces language confusion issues, especially in mismatched domain scenarios. In this paper, we decouple language confusion in LR-MoE into confusion in self-attention and router. To alleviate the language confusion in self-attention, based on LR-MoE, we propose to apply attention-MoE architecture for MASR. In our new architecture, MoE is utilized not only on feedforward network (FFN) but also on self-attention. In addition, to improve the robustness of the LID-based router on language confusion, we propose expert pruning and router augmentation methods. Combining the above, we get the boosted language-routing MoE (BLR-MoE) architecture. We verify the effectiveness of the proposed BLR-MoE in a 10,000-hour MASR dataset.
Lifeng Zhou 0003, Yuke Li 0002, Binbin Du
ICASSP5
2025 Bridging the Modality Gap for Speech-image Retrieval with Text Supervision
abstract
In recent years, while the performance of speech-image retrieval has improved significantly, it still lags behind that of image-text retrieval. Leveraging the text modality to enhance speech-image retrieval remains a promising research direction. In this paper, we propose to leverage text supervision to facilitate the alignment between speech and image feature spaces via an automatic speech recognition (ASR) auxiliary task. Specifically, our model is trained with a multi-task learning framework, which combines ASR and speech-image contrastive learning tasks. On this basis, benefiting from the ASR module to obtain text modality information, we introduce an ensemble mechanism to further enhance speech-image retrieval performance. Experimental results on two publicly available datasets Flickr8k and SpokenCOCO indicate that the introduction of the ASR task improves the mean R@1 by 2.7% and 2%, respectively, compared to the previous state-of-the-art (SOTA) method. Furthermore, the results demonstrate that the ensemble mechanism further enhances the performance of speech-image retrieval.
Lifeng Zhou 0003, Yuke Li 0002
ICASSP3
2025 SELECT: Detecting Label Errors in Real-world Scene Text Data
abstract
We introduce SELECT (Scene tExt Label Errors deteCTion), a novel approach that leverages multi-modal training to detect label errors in real-world scene text datasets. Utilizing an image-text encoder and a character-level tokenizer, SELECT addresses the issues of variable-length sequence labels, label sequence misalignment, and character-level errors, outperforming existing methods in accuracy and practical utility. In addition, we introduce Similarity-based Sequence Label Corruption (SSLC), a process that intentionally introduces errors into the training labels to mimic real-world error scenarios during training. SSLC not only can cause a change in the sequence length but also takes into account the visual similarity between characters during corruption. Our method is the first to detect label errors in real-world scene text datasets successfully accounting for variable-length labels. Experimental results demonstrate the effectiveness of SELECT in detecting label errors and improving STR accuracy on real-world text datasets, showcasing its practical utility.
Yifeng Hu, Yuke Li 0002
MMAsia4
2024 Differentiable Resolution Compression and Alignment for Efficient Video Classification and Retrieval
abstract
Optimizing video inference efficiency has become increasingly important with the growing demand for video analysis in various fields. Some existing methods achieve high efficiency by explicit discard of spatial or temporal information, which poses challenges in fast-changing and fine-grained scenarios. To address these issues, we propose an efficient video representation network with Differentiable Resolution Compression and Alignment, which compresses non-essential information in the early stage of the network to reduce computational costs while maintaining consistent temporal correlations. Specifically, we leverage a Differentiable Context-aware Compression Module to encode the saliency and non-saliency frame features, refining and updating the features into a high-low resolution video sequence. To process the new sequence, we introduce a new Resolution-Align Transformer Layer to capture global temporal correlations among frame features with different resolutions, while reducing spatial computation costs quadratically by utilizing fewer spatial tokens in low-resolution non-saliency frames. The entire network can be end-to-end optimized via the integration of the differentiable compression module. Experimental results show that our method achieves the best trade-off between efficiency and performance on near-duplicate video retrieval and competitive results on dynamic video classification compared to state-of-the-art methods.Code:https://github.com/dun-research/DRCA
Yuke Li 0002, Haoran Fu
ICASSP3
2024 HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition
abstract
Action recognition in videos poses a challenge due to its high computational cost, especially for Joint Space-Time video transformers (Joint VT). Despite their effectiveness, the excessive number of tokens in such architectures significantly limits their efficiency. In this paper, we propose HaltingVT, an efficient video transformer adaptively removing redundant video patch tokens, which is primarily composed of a Joint VT and a Glimpser module. Specifically, HaltingVT applies data-adaptive token reduction at each layer, resulting in a significant reduction in the overall computational cost. Besides, the Glimpser module quickly removes redundant tokens in shallow transformer layers, which may even be misleading for video recognition tasks based on our observations. To further encourage HaltingVT to focus on the key motion-related information in videos, we design an effective Motion Loss during training. HaltingVT acquires video analysis capabilities and token halting compression strategies simultaneously in a unified training process, without requiring additional training procedures or sub-networks. On the Mini-Kinetics dataset [1], we achieved 75.0% top-1 ACC with 24.2 GFLOPs, as well as 67.2% top-1 ACC with an extremely low 9.9 GFLOPs. The code is available at https://github.com/dun-research/HaltingVT.
Ruoxuan Cui, Yuke Li 0002, Haoqi Zhu
ICASSP3
2024 DeCMG: Denoise with Cross-modality Guidance Makes Better Text-Video Retrieval
abstract
The effectiveness of text-video retrieval depends on establishing precise semantic correspondence between text and video. However, current CLIP-based [1] approaches, have encountered challenges in adequately learning temporal and cross-modality features due to the lack of prior knowledge and the scarcity of labeled training data. To tackle this issue, we propose DeCMG (Denoise with Cross-Modality Guidance), a novel method that enhances fine-grained text-video alignment through cross-modality collaboration. The collaboration in DeCMG allows the features in one modality to be better optimized by performing a denoising task in another modality with a guidance vector. Specifically, DeCMG incorporates two self-generated denoising tasks (TextDe and VisualDe) to achieve alignment at the word-video and frame-text levels, respectively. These tasks are implemented as plugins, resulting in only a minor increase in training time and can be entirely removed during inference. Across four benchmark datasets, DeCMG yielded average improvements of 1.1% in text-to-video and 1.8% in video-to-text Recall@1, surpassing our baseline models (X-CLIP [2] and X-Pool [3]). These experimental results validate the efficiency and effectiveness of DeCMG.
Yuke Li 0002
ICME2
2024 Coarse-to-fine Alignment Makes Better Speech-image Retrieval
abstract
In this paper, we propose a novel framework for speech-image retrieval. We utilize speech-image contrastive (SIC) learning tasks to align speech and image representations at a coarse level and speech-image matching (SIM) learning tasks to further refine the fine-grained cross-modal alignment. SIC and SIM learning tasks are jointly trained in a unified manner. To optimize the learning process, we utilize an embedding queue that facilitates efficient sampling of high-quality and diverse negative representations during SIC learning. Additionally, it enhances the learning of SIM tasks by effectively mining hard negatives based on contrastive similarities calculated in SIC tasks. To further optimize learning under noisy supervision, we incorporate momentum distillation into the training process. Experimental results show that our framework outperforms the state-of-the-art method by more than 4% in R@1 on two benchmark datasets for the speech-image retrieval tasks. Moreover, as observed in zero-shot experiments, our framework demonstrates excellent generalization capabilities.
Lifeng Zhou 0003, Yuke Li 0002
ICME2
2024 Learning from Back Chunks: Acquiring More Future Knowledge for Streaming ASR Models via Self Distillation
Yuke Li 0002, Binbin Du, Haoqi Zhu, Liang Ruan
INTERSPEECH3
2024 Cross-Modal Denoising: A Novel Training Paradigm for Enhancing Speech-Image Retrieval
abstract
The success of speech-image retrieval relies on establishing an effective alignment between speech and image.Existing methods often model cross-modal interaction through simple cosine similarity of the global feature of each modality, which fall short in capturing fine-grained details within modalities.To address this issue, we introduce an effective framework and a novel learning task named cross-modal denoising (CMD) to enhance cross-modal interaction to achieve finerlevel cross-modal alignment.Specifically, CMD is a denoising task designed to reconstruct semantic features from noisy features within one modality by interacting features from another modality.Notably, CMD operates exclusively during model training and can be removed during inference without adding extra inference time.The experimental results demonstrate that our framework outperforms the state-of-the-art method by 2.0% in mean R@1 on the Flickr8k dataset and by 1.7% in mean R@1 on the SpokenCOCO dataset for the speech-image retrieval tasks, respectively.These experimental results validate the efficiency and effectiveness of our framework.
Lifeng Zhou 0003, Yuke Li 0002, Haoqi Zhu
INTERSPEECH2
2024 Enhancing Unified Streaming and Non-Streaming ASR Through Curriculum Learning With Easy-To-Hard Tasks
abstract
We expect a unified ASR model to deliver high performance in both streaming and non-streaming modes. However, a core challenge is that the lack of global contextual information in streaming ASR inherently hinders its performance from matching the non-streaming counterpart. Drawing inspiration from the human learning manner from easy concepts to difficult ones, we introduce a curriculum learning framework to enhance the training of unified ASR models. This framework strategically increases task complexity in a graduated, easy-to-hard order. Specifically, we develop a structured curriculum that begins with an elementary course focused on training a non-streaming model, progresses to an intermediate course for training an initial unified ASR model, and culminates in an advanced course designed to mutual promotion between these two modes via contrastive training. Experimental results on AISHELL-1 and AISHELL-2 show that our method achieves significant improvements in two modes.
Yuke Li 0002, Lifeng Zhou 0003, Binbin Du, Haoqi Zhu
SLT2
2023 LAE-ST-MOE: Boosted Language-Aware Encoder Using Speech Translation Auxiliary Task for E2E Code-Switching ASR
abstract
Recently, to mitigate the confusion between different languages in code-switching (CS) automatic speech recognition (ASR), the conditionally factorized models, such as the language-aware encoder (LAE), explicitly disregard the contextual information between different languages. However, this information may be helpful for ASR modeling. To alleviate this issue, we propose the LAE-ST-MoE framework. It incorporates speech translation (ST) tasks into LAE and utilizes ST to learn the contextual information between different languages. It introduces a task-based mixture of expert modules, employing separate feed-forward networks for the ASR and ST tasks. Experimental results on the ASRU 2019 Mandarin-English CS challenge dataset demonstrate that, compared to the LAE-based CTC, the LAE-ST-MoE model achieves a 9.26 % mix error reduction on the CS test with the same decoding parameter. Moreover, the well-trained LAE-ST-MoE model can perform ST tasks from CS speech to Mandarin or English text.
Yuke Li 0002, Binbin Du, Haoran Fu
ASRU3
2023 Improving CTC-Based ASR Models With Gated Interlayer Collaboration
abstract
The CTC-based automatic speech recognition (ASR) models without the external language model usually lack the capacity to model conditional dependencies and textual interactions. In this paper, we present a Gated Interlayer Collaboration (GIC) mechanism to improve the performance of CTC-based models, which introduces textual information into the model and thus relaxes the conditional independence assumption of CTC-based models. Specifically, we consider the weighted sum of token embeddings as the textual representation for each position, where the position-specific weights are the softmax probability distribution constructed via inter-layer auxiliary CTC losses. The textual representations are then fused with acoustic features by developing a gate unit. Experiments on AISHELL-1 [1], TEDLIUM2 [2], and AI-DATATANG [3] corpora show that the proposed method outperforms several strong baselines.
Yuke Li 0002, Binbin Du
ICASSP2
2023 3D-CSL: Self-Supervised 3D Context Similarity Learning for Near-Duplicate Video Retrieval
abstract
In this paper, we introduce 3D-CSL, a compact pipeline for Near-Duplicate Video Retrieval (NDVR), and explore a novel self-supervised learning strategy for video similarity learning. Most previous NDVR methods depend a lot on pair-wise labeled data, so that be limited by the scale of datasets and cannot optimize complex but efficient backbones, e.g., 3D transformers. In order to break this limitation, we explore the self-supervised similarity learning for the NDVR task and propose FCS loss, a novel triplet loss, and ShotMix, a novel video-specific augmentation, which enhances the self-supervised video similarity learning significantly. With this premise, the compact 3D pipeline we propose shows a great advantage in extracting global spatiotemporal dependencies in videos and achieves the best balance between efficiency and effectiveness. Furthermore, we also propose PredMAE to pretrain the 3D transformer with video prediction task as a pretext task to boost the downstream NDVR task without any human labels. The experiments on FIVR-200K and CC_WEB_VIDEO demonstrate the superiority and reliability of our method, which achieves the state-of-the-art performance on clip-level NDVR. Code is released in https://github.com/dun-research/3D-CSL
Yuke Li 0002
ICIP3
2023 Language-Routing Mixture of Experts for Multilingual and Code-Switching Speech Recognition
Yuke Li 0002, Binbin Du
INTERSPEECH3
2023 Enhancing the Unified Streaming and Non-streaming Model with Contrastive Learning
Yuke Li 0002, Binbin Du
INTERSPEECH2