Mingshuang Luo

dblp:260/0750 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-4287-9055ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%
Artificial intelligence
2 papers
Generative modeling · 51% Language models and text generation · 31% 3D vision · 18%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.912025
Morph: a Motion-Free Physics Optimization Framework for Human Motion Generation · ICCV 2025
Machine learning › Generative modeling
motion generation
0.812024
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation · NeurIPS 2024
Computer vision › 3D vision
physical simulation
0.312025
Morph: a Motion-Free Physics Optimization Framework for Human Motion Generation · ICCV 2025
Natural language and speech › Language models and text generation
large language model
0.212024
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation · NeurIPS 2024
Natural language and speech › Language models and text generation
multimodal language model
0.212024
M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

reward shaping · 1.7reinforcement learning · 1.7physics simulator · 1.7large language model · 0.8discrete vector quantization · 0.8
YearPublicationVenuePosition
2025 Morph: a Motion-Free Physics Optimization Framework for Human Motion Generation
abstract
Human motion generation has been widely studied due to its crucial role in areas such as digital humans and humanoid robot control. However, many current motion generation approaches disregard physics constraints, frequently resulting in physically implausible motions with pronounced artifacts such as floating and foot sliding. Meanwhile, training an effective motion physics optimizer with noisy motion data remains largely unexplored. In this paper, we propose \textbf{Morph}, a \textbf{Mo}tion-F\textbf{r}ee \textbf{ph}ysics optimization framework, consisting of a Motion Generator and a Motion Physics Refinement module, for enhancing physical plausibility without relying on expensive real-world motion data. Specifically, the motion generator is responsible for providing large-scale synthetic, noisy motion data, while the motion physics refinement module utilizes these synthetic data to learn a motion imitator within a physics simulator, enforcing physical constraints to project the noisy motions into a physically-plausible space. Additionally, we introduce a prior reward module to enhance the stability of the physics optimization process and generate smoother and more stable motions. These physically refined motions are then used to fine-tune the motion generator, further enhancing its capability. This collaborative training paradigm enables mutual enhancement between the motion generator and the motion physics refinement module, significantly improving practicality and robustness in real-world applications. Experiments on both text-to-motion and music-to-dance generation tasks demonstrate that our framework achieves state-of-the-art motion quality while improving physical plausibility drastically. Project page: https://interestingzhuo.github.io/Morph-Page/.
Mingshuang Luo, Ruibing Hou, Zimo Liu
ICCV2
2024 M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
abstract
This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities. We employ discrete vector quantization for multimodal conditional signals, such as text, music and motion/dance, enabling seamless integration into a large language model (LLM) with a single vocabulary. The second involves modeling motion generation directly in the raw motion space. This strategy circumvents the information loss associated with a discrete tokenizer, resulting in more detailed and comprehensive motion generation. Third, M$^3$GPT learns to model the connections and synergies among various motion-relevant tasks. Text, the most familiar and well-understood modality for LLMs, is utilized as a bridge to establish connections between different motion tasks, facilitating mutual reinforcement. To our knowledge, M$^3$GPT is the first model capable of comprehending and generating motions based on multiple signals. Extensive experiments highlight M$^3$GPT's superior performance across various motion-relevant tasks and its powerful zero-shot generalization capabilities for extremely challenging tasks. Project page: \url{https://github.com/luomingshuang/M3GPT}.
Mingshuang Luo, Ruibing Hou, Hong Chang 0001, Zimo Liu, Shiguang Shan
NeurIPS1
2024 CELA-MFP: a contrast-enhanced and label-adaptive framework for multi-functional therapeutic peptides prediction
abstract
Functional peptides play crucial roles in various biological processes and hold significant potential in many fields such as drug discovery and biotechnology. Accurately predicting the functions of peptides is essential for understanding their diverse effects and designing peptide-based therapeutics. Here, we propose CELA-MFP, a deep learning framework that incorporates feature Contrastive Enhancement and Label Adaptation for predicting Multi-Functional therapeutic Peptides. CELA-MFP utilizes a protein language model (pLM) to extract features from peptide sequences, which are then fed into a Transformer decoder for function prediction, effectively modeling correlations between different functions. To enhance the representation of each peptide sequence, contrastive learning is employed during training. Experimental results demonstrate that CELA-MFP outperforms state-of-the-art methods on most evaluation metrics for two widely used datasets, MFBP and MFTP. The interpretability of CELA-MFP is demonstrated by visualizing attention patterns in pLM and Transformer decoder. Finally, a user-friendly online server for predicting multi-functional peptides is established as the implementation of the proposed CELA-MFP and can be freely accessed at http://dreamai.cmii.online/CELA-MFP.
Yitian Fang, Mingshuang Luo, Zhixiang Ren, Leyi Wei
Briefings Bioinform.2
2023 Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
abstract
Knowledge distillation (KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour of a teacher model. However, traditional KD methods suffer from teacher label storage issue, especially when the training corpora are large. Although on-the-fly teacher label generation tackles this issue, the training speed is significantly slower as the teacher model has to be evaluated every batch. In this paper, we reformulate the generation of teacher label as a codec problem. We propose a novel Multi-codebook Vector Quantization (MVQ) approach that compresses teacher embeddings to codebook indexes (CI). Based on this, a KD training framework (MVQ-KD) is proposed where a student model predicts the CI generated from the embeddings of a self-supervised pre-trained teacher model. Experiments on the LibriSpeech clean-100 hour show that MVQ-KD framework achieves comparable performance as traditional KD methods (11, 12), while requiring 256 times less storage. When the full LibriSpeech dataset is used, MVQ-KD framework results in 13.8% and 8.2% relative word error rate reductions (WERRs) for non -streaming transducer on test-clean and test-other and 4.0% and 4.9% for streaming transducer. The implementation of this work is already released as a part of the open-source project icefall1.
Liyong Guo, Xiaoyu Yang 0005, Quandong Wang, Yuxiang Kong, Zengwei Yao, Fan Cui, Wei Kang 0006, Long Lin, Mingshuang Luo, Piotr Zelasko, Daniel Povey
ICASSP10
2023 Fast and Parallel Decoding for Transducer
abstract
The transducer architecture is becoming increasingly popular in the field of speech recognition, because it is naturally streaming as well as high in accuracy. One of the drawbacks of transducer is that it is difficult to decode in a fast and parallel way due to an unconstrained number of symbols that can be emitted per time step.In this work, we introduce a constrained version of transducer loss to learn strictly monotonic alignments between the sequences; we also improve the standard greedy search and beam search algorithms by limiting the number of symbols that can be emitted per time step in transducer decoding, making it more efficient to decode in parallel with batches. Furthermore, we propose an finite state automaton-based (FSA) parallel beam search algorithm that can run with graphs on GPU efficiently. The experiment results show that we achieve slight word error rate (WER) improvement as well as significant speedup in decoding. Our work is open-sourced and publicly available1.
Wei Kang 0006, Liyong Guo, Long Lin, Mingshuang Luo, Zengwei Yao, Xiaoyu Yang 0005, Piotr Zelasko, Daniel Povey
ICASSP5
2022 Pruned RNN-T for fast, memory-efficient ASR training
abstract
The RNN-Transducer (RNN-T) framework for speech recognition has been growing in popularity, particularly for deployed real-time ASR systems, because it combines high accuracy with naturally streaming recognition.One of the drawbacks of RNN-T is that its loss function is relatively slow to compute, and can use a lot of memory.Excessive GPU memory usage can make it impractical to use RNN-T loss in cases where the vocabulary size is large: for example, for Chinese character-based ASR.We introduce a method for faster and more memoryefficient RNN-T loss computation.We first obtain pruning bounds for the RNN-T recursion using a simple joiner network that is linear in the encoder and decoder embeddings; we can evaluate this without using much memory.We then use those pruning bounds to evaluate the full, non-linear joiner network.The code is open-sourced and publicly available.
Liyong Guo, Wei Kang 0006, Long Lin, Mingshuang Luo, Zengwei Yao, Daniel Povey
INTERSPEECH5
2020 Synchronous Bidirectional Learning for Multilingual Lip Reading
Mingshuang Luo, Xilin Chen 0001, Shiguang Shan
BMVC1
2020 Pseudo-Convolutional Policy Gradient for Sequence-to-Sequence Lip-Reading
abstract
Lip-reading aims to infer the speech content from the lip movement sequence and can be seen as a typical sequence-to-sequence (seq2seq) problem which translates the input image sequence of lip movements to the text sequence of the speech content. However, the traditional learning process of seq2seq models always suffers from two problems: the exposure bias resulted from the strategy of “teacher-forcing”, and the inconsistency between the discriminative optimization target (usually the cross-entropy loss) and the final evaluation metric (usually the character/word error rate). In this paper, we propose a novel pseudo-convolutional policy gradient (PCPG) based method to address these two problems. On the one hand, we introduce the evaluation metric (refers to the character error rate in this paper) as a form of reward to optimize the model together with the original discriminative target. On the other hand, inspired by the local perception property of convolutional operation, we perform a pseudo-convolutional operation on the reward and loss dimension, so as to take more context around each time step into account to generate a robust reward and loss for the whole optimization. Finally, we perform a thorough comparison and evaluation on both the word-level and sentence-level benchmarks. The results show a significant improvement over other related methods, and report either a new state-of-the-art performance or a competitive accuracy on all these challenging benchmarks, which clearly proves the advantages of our approach.
Mingshuang Luo, Shiguang Shan, Xilin Chen 0001
FG1