VLDB 2026 Research / reviewers in the wild / expert
Ming Sun 0013
dblp:39/1471-13
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Long-Form Fuzzy Speech-to-Text Alignment for 1000+ LanguagesabstractConventional speech-to-text forced alignment typically operates at the utterance level. In practice, however, we do not usually have short segments (e.g., 10 seconds) of audio with exact, verbatim transcriptions (e.g., the LibriSpeech corpus) as in lab conditions. Instead, audio often comes in long-form (e.g., an hour-long lecture recording), and the available transcription may be non-verbatim or include unspoken annotations, making it misaligned with the actual speech. This motivates the need for long-form fuzzy speech-to-text alignment, which has practical applications - for example, preparing segmented supervised audio data for training machine learning models. We demonstrate the Torchaudio long-form aligner, which supports such use cases. Moreover, it can be equipped with any CTC model that predicts frame-wise labels, turning the model into a robust and powerful aligner. Ruizhe Huang, Xiaohui Zhang 0007, Zhaoheng Ni, Moto Hira, Jeff Hwang, Vineel Pratap, Ju Lin, Ming Sun 0013, Florian Metze |
ASRU | 8 |
| 2025 | MMW: Side Talk Rejection Multi-Microphone Whisper On Smart GlassesabstractSmart glasses are increasingly positioned as the nextgeneration interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in real-world noisy environments remains a major challenge, particularly due to interference from side speech. In this work, we introduce a novel side-talk rejection multimicrophone Whisper (MMW) framework for smart glasses, incorporating three key innovations. First, we propose a Mix Block based on a Tri-Mamba architecture to effectively fuse multi-channel audio at the raw waveform level, while maintaining compatibility with streaming processing. Second, we design a Frame Diarization Mamba Layer to enhance frame-level side-talk suppression, facilitating more efficient fine-tuning of Whisper models. Third, we employ a Multi-Scale Group Relative Policy Optimization (GRPO) strategy to jointly optimize frame-level and utterance-level side speech suppression. Experimental evaluations demonstrate that the proposed MMW system can reduce the word error rate (WER) by 4.95% in noisy conditions. Yiteng Huang, Yangyang Shi, Saurabh Adya, Ming Sun 0013, Florian Metze |
ASRU | 7 |
| 2025 | Directional Source Separation for Robust Speech Recognition on Smart GlassesabstractModern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this work investigates directional source separation using the multi-microphone array. We explore multiple beamformers to assist source separation by strengthening the directional properties of speech signals. In addition to relying on predetermined beamformers, we investigate neural beamforming in multi-channel source separation, demonstrating that automatic learning directional characteristics effectively improves separation quality. Furthermore, we investigate the training strategies for ASR when utilizing separated outputs. Our results suggest that jointly training a directional speech separation and ASR model achieves the best overall performance while balancing the wearer and conversation partner’s performance. Tiantian Feng, Ju Lin, Yiteng Huang, Weipeng He, Kaustubh Kalgaonkar, Niko Moritz, Ming Sun 0013, Frank Seide |
ICASSP | 9 |
| 2025 | Effective Integration of KAN for Keyword SpottingabstractKeyword spotting (KWS) is an important speech processing component for smart devices with voice assistance capability. In this paper, we investigate if Kolmogorov-Arnold Networks (KAN) can be used to enhance the performance of KWS. We explore various approaches to integrate KAN for a model architecture based on 1D Convolutional Neural Networks (CNN). We find that KAN is effective at modeling high-level features in lower-dimensional spaces, resulting in improved KWS performance when integrated appropriately. The findings shed light on understanding KAN for speech processing tasks and on other modalities for future researchers. Anfeng Xu, Biqiao Zhang, Shuyu Kong, Yiteng Huang, Sangeeta Srivastava, Ming Sun 0013 |
ICASSP | 7 |
| 2025 | Directional Speech Recognition with Full-Duplex Capability
Ju Lin, Yiteng Huang, Ming Sun 0013, Frank Seide, Florian Metze |
INTERSPEECH | 3 |
| 2025 | MASV: Speaker Verification with Global and Local Context Mamba
Yiteng Huang, Ming Sun 0013, Xinhao Mei, Yangyang Shi, Florian Metze |
INTERSPEECH | 4 |
| 2025 | Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Prashant Rawat, Sangeeta Srivastava, Ming Sun 0013, Florian Metze |
INTERSPEECH | 10 |
| 2024 | AGADIR: Towards Array-Geometry Agnostic Directional Speech RecognitionabstractWearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise.When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses.This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion. Ju Lin, Niko Moritz, Yiteng Huang, Ruiming Xie, Ming Sun 0013, Christian Fügen, Frank Seide |
ICASSP | 5 |
| 2024 | Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
Shuyu Kong, Biqiao Zhang, Yiteng Huang, Mumin Jin, Ming Sun 0013 |
INTERSPEECH | 7 |
| 2023 | Disentangled Training with Adversarial Examples for Robust Small-Footprint Keyword SpottingabstractA keyword spotting (KWS) engine continuously running on the device is exposed to various speech signals that are usually unseen beforehand. It is a challenging problem to build a small-footprint and high-performing KWS model with robustness under different acoustic environments. In this paper, we explore how to effectively apply adversarial examples to improve KWS robustness. We propose datasource-aware disentangled learning with adversarial examples to reduce the mismatch between the original and adversarial data as well as the mismatch across original training datasources. The KWS model architecture is based on depth-wise separable convolution and a simple attention module. Experimental results demonstrate that the proposed learning strategy improves false reject rate by 40.31% at 1% false accept rate on the internal dataset, compared to the strongest baseline without adversarial examples. Our best-performing system achieves 98.06% accuracy on the Google Speech Commands V1 dataset. Biqiao Zhang, Yiteng Huang, Shang-Wen Li 0001, Ming Sun 0013 |
ICASSP | 6 |
| 2023 | Handling the Alignment for Wake Word Detection: A Comparison Between Alignment-Based, Alignment-Free and Hybrid Approaches
Vinicius Ribeiro, Yiteng Huang, Yuan Shangguan, Ming Sun 0013 |
INTERSPEECH | 6 |