VLDB 2026 Research / reviewers in the wild / expert
Liming Wang 0003
dblp:51/8-3
· DBLP profile ↗
14ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-5562-8386ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards unsupervised speech recognition without pronunciation models
Junrui Ni, Liming Wang 0003, Yang Zhang 0001, Kaizhi Qian, Heting Gao, Mark Hasegawa-Johnson, James R. Glass, Chang Dong Yoo |
Speech Commun. | 2 |
| 2025 | Recognizing Dementia from Neuropsychological Tests with State Space ModelsabstractEarly detection of dementia is critical for timely medical intervention and improved patient outcomes. Neuropsychological tests are widely used for cognitive assessment but have traditionally relied on manual scoring. Automatic dementia classification (ADC) systems aim to infer cognitive decline directly from speech recordings of such tests. We propose Demenba, a novel ADC framework based on state space models, which scale linearly in memory and computation with sequence length. Trained on over 1,000 hours of cognitive assessments administered to Framingham Heart Study participants, some of whom were diagnosed with dementia through adjudicated review, our method outperforms prior approaches in fine-grained dementia classification by 21%, while using fewer parameters. We further analyze its scaling behavior and demonstrate that our model gains additional improvement when fused with large language models, paving the way for more transparent and scalable dementia assessment tools12.1Code: https://github.com/lwang114/Demenba2This work was supported by the Framingham Heart Study’s National Heart, Lung, and Blood Institute contract N01-HC-25195; National Institutes of Health grants U19-AG068753, R01- AG016495, R01-AG008122, R01AG033040. The authors would also like to thank the staff and participants of the Framingham Heart Study. Liming Wang 0003, Saurabhchand Bhati, Cody Karjadi, Rhoda Au, James R. Glass |
ASRU | 1 |
| 2025 | Can Diffusion Models Disentangle? A Theoretical PerspectiveabstractThis paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations with commonly used weak supervision such as partial labels and multiple views. Within this framework, we establish identifiability conditions for diffusion models to disentangle latent variable models with \emph{stochastic}, \emph{non-invertible} mixing processes. We also prove \emph{finite-sample global convergence} for diffusion models to disentangle independent subspace models. To validate our theory, we conduct extensive disentanglement experiments on subspace recovery in latent subspace Gaussian mixture models, image colorization, denoising, and voice conversion for speech classification. Our experiments show that training strategies inspired by our theory, such as style guidance regularization, consistently enhance disentanglement performance. Liming Wang 0003, Muhammad Jehanzeb Mirza, Yishu Gong, Yuan Gong 0001, Jiaqi Zhang 0006, Brian Tracey, Katerina Placek, Marco Vilela, James R. Glass |
NeurIPS | 1 |
| 2024 | Unsupervised Speech Recognition with N-skipgram and Positional Unigram MatchingabstractTraining unsupervised speech recognition systems presents challenges due to GAN-associated instability, misalignment between speech and text, and significant memory demands. To tackle these challenges, we introduce a novel ASR system, ESPUM. This system harnesses the power of lower-order N-skipgrams (up to N = 3) combined with positional unigram statistics gathered from a small batch of samples. Evaluated on the TIMIT benchmark, our model showcases competitive performance in ASR and phoneme segmentation tasks. Access our publicly available code at https://github.com/lwang114/GraphUnsupASR. Liming Wang 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 1 |
| 2024 | Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
Liming Wang 0003, Yuan Gong 0001, Nauman Dawalatabad, Marco Vilela, Katerina Placek, Brian Tracey, Yishu Gong, Alan Premasiri, Fernando Vieira, James R. Glass |
INTERSPEECH | 1 |
| 2023 | A Theory of Unsupervised Speech RecognitionabstractUnsupervised speech recognition (ASR-U) is the problem of learning automatic speech recognition (ASR) systems from unpaired speech-only and text-only corpora.While various algorithms exist to solve this problem, a theoretical framework is missing to study their properties and address such issues as sensitivity to hyperparameters and training instability.In this paper, we proposed a general theoretical framework to study the properties of ASR-U systems based on random matrix theory and the theory of neural tangent kernels.Such a framework allows us to prove various learnability conditions and sample complexity bounds of ASR-U.Extensive ASR-U experiments on synthetic languages with three classes of transition graphs provide strong empirical evidence for our theory (code available at cactuswith- thoughts/UnsupASRTheory.git). Liming Wang 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ACL (1) | 1 |
| 2022 | Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech RecognitionabstractPhonemes are defined by their relationship to words: changing a phoneme changes the word.Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to underresourced speech technology.In this paper, we bridge the gap between the linguistic and statistical definition of phonemes and propose a novel neural discrete representation learning model for self-supervised learning of phoneme inventory with raw speech and word labels.Given the availability of phoneme segmentation and some mild conditions, we prove that the phoneme inventory learned by our approach converges to the true one with an exponentially low error rate.Moreover, in experiments on TIMIT and Mboshi benchmarks, our approach consistently learns a better phonemelevel representation and achieves a lower error rate in a zero-resource phoneme recognition task than previous state-of-the-art selfsupervised representation learning algorithms. Liming Wang 0003, Siyuan Feng 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ACL (1) | 1 |
| 2022 | Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition
Junrui Ni, Liming Wang 0003, Heting Gao, Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2021 | Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image RetrievalabstractMultimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conducting retrieval and word discovery experiments on MSCOCO and Flickr8k, and empirically demonstrate that both neural MT with self-attention and statistical MT achieve word discovery scores that are superior to those of a state-of-the-art neural retrieval system, outperforming it by 2% and 5% alignment F1 scores respectively. Liming Wang 0003, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 1 |
| 2020 | A DNN-HMM-DNN Hybrid Model for Discovering Word-Like Units from Spoken Captions and Image Regions
Liming Wang 0003, Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2020 | Speech Technology for Unwritten LanguagesabstractSpeech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible. Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 10 |
| 2020 | Multimodal Word Discovery and Retrieval With Spoken Descriptions and Visual ConceptsabstractIn the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this article, four multimodal word discovery systems are demonstrated: three models based on statistical machine translation (SMT) and one based on neural machine translation (NMT). The systems are trained with phonetic transcriptions, MFCC and multilingual bottleneck features (MBN). On the phone-level, the SMT outperforms the NMT model, achieving a 61.6% F1 score in the phone-level word discovery task on Flickr30k. On the audio-level, we compared our models with the existing ES-KMeans algorithm for word discovery and present some of the challenges in multimodal spoken word discovery. Liming Wang 0003, Mark Hasegawa-Johnson |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Multimodal Word Discovery and Retrieval with Phone Sequence and Image Concepts
Liming Wang 0003, Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2018 | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 WorkshopabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech. Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux |
ICASSP | 18 |