VLDB 2026 Research / reviewers in the wild / expert
Ju Lin
dblp:193/6454
· DBLP profile ↗
21ranked-venue papers
10as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Long-Form Fuzzy Speech-to-Text Alignment for 1000+ LanguagesabstractConventional speech-to-text forced alignment typically operates at the utterance level. In practice, however, we do not usually have short segments (e.g., 10 seconds) of audio with exact, verbatim transcriptions (e.g., the LibriSpeech corpus) as in lab conditions. Instead, audio often comes in long-form (e.g., an hour-long lecture recording), and the available transcription may be non-verbatim or include unspoken annotations, making it misaligned with the actual speech. This motivates the need for long-form fuzzy speech-to-text alignment, which has practical applications - for example, preparing segmented supervised audio data for training machine learning models. We demonstrate the Torchaudio long-form aligner, which supports such use cases. Moreover, it can be equipped with any CTC model that predicts frame-wise labels, turning the model into a robust and powerful aligner. Ruizhe Huang, Xiaohui Zhang 0007, Zhaoheng Ni, Moto Hira, Jeff Hwang, Vineel Pratap, Ju Lin, Ming Sun 0013, Florian Metze |
ASRU | 7 |
| 2025 | Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
Guilin Zhang, Wulan Guo, Srinivas Vippagunta, Suchitra Raman, Shreeshankar Chatterjee, Ju Lin, Mary Schladenhauffen, Jeffrey Luo, Hailong Jiang |
IEEE Big Data | 7 |
| 2025 | Directional Source Separation for Robust Speech Recognition on Smart GlassesabstractModern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this work investigates directional source separation using the multi-microphone array. We explore multiple beamformers to assist source separation by strengthening the directional properties of speech signals. In addition to relying on predetermined beamformers, we investigate neural beamforming in multi-channel source separation, demonstrating that automatic learning directional characteristics effectively improves separation quality. Furthermore, we investigate the training strategies for ASR when utilizing separated outputs. Our results suggest that jointly training a directional speech separation and ASR model achieves the best overall performance while balancing the wearer and conversation partner’s performance. Tiantian Feng, Ju Lin, Yiteng Huang, Weipeng He, Kaustubh Kalgaonkar, Niko Moritz, Ming Sun 0013, Frank Seide |
ICASSP | 2 |
| 2025 | M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart GlassesabstractThe growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently trained models, which may not benefit from large amounts of unlabeled data. In this paper, we propose M-BEST-RQ, the first multi-channel speech foundation model for smart glasses, which is designed to leverage large-scale self-supervised learning (SSL) in an array-geometry agnostic approach. While prior work on multi-channel speech SSL only evaluated on simulated settings, we curate a suite of real downstream tasks to evaluate our model, namely (i) conversational automatic speech recognition (ASR), (ii) spherical active source localization, and (iii) glasses wearer voice activity detection, which are sourced from the MMCSG and EasyCom datasets. We show that a general-purpose M-BEST-RQ encoder is able to match or surpass supervised models across all tasks. For the conversational ASR task in particular, using only 8 hours of labeled speech, our model outperforms a supervised ASR baseline that is trained on 2000 hours of labeled data, which demonstrates the effectiveness of our approach. Desh Raj, Ju Lin, Niko Moritz, Junteng Jia, Gil Keren, Egor Lakomkin, Yiteng Huang, Jacob Donley, Jay Mahadeokar, Ozlem Kalinli |
ICASSP | 3 |
| 2025 | Directional Speech Recognition with Full-Duplex Capability
Ju Lin, Yiteng Huang, Ming Sun 0013, Frank Seide, Florian Metze |
INTERSPEECH | 1 |
| 2025 | Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Prashant Rawat, Sangeeta Srivastava, Ming Sun 0013, Florian Metze |
INTERSPEECH | 2 |
| 2024 | AGADIR: Towards Array-Geometry Agnostic Directional Speech RecognitionabstractWearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise.When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses.This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion. Ju Lin, Niko Moritz, Yiteng Huang, Ruiming Xie, Ming Sun 0013, Christian Fügen, Frank Seide |
ICASSP | 1 |
| 2024 | Exploring the Feasibility of Automated Data Standardization using Large Language Models for Seamless PositioningabstractWe propose a feasibility study for real-time automated data standardization leveraging Large Language Models (LLMs) to enhance seamless positioning systems in IoT environments. By integrating and standardizing heterogeneous sensor data from smartphones, IoT devices, and dedicated systems such as Ultra-Wideband (UWB), our study ensures data compatibility and improves positioning accuracy using the Extended Kalman Filter (EKF). The core components include the Intelligent Data Standardization Module (IDSM), which employs a fine-tuned LLM to convert varied sensor data into a standardized format, and the Transformation Rule Generation Module (TRGM), which automates the creation of transformation rules and scripts for ongoing data standardization. Evaluated in real-time environments, our study demonstrates adaptability and scalability, enhancing operational efficiency and accuracy in seamless navigation. This study underscores the potential of advanced LLMs in overcoming sensor data integration complexities, paving the way for more scalable and precise IoT navigation solutions. Max Jwo Lem Lee, Ju Lin, Li-Ta Hsu |
IPIN | 2 |
| 2023 | Egocentric Audio-Visual Noise SuppressionabstractThis paper studies audio-visual noise suppression for egocentric videos -where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker’s view of the outside world. This setting is different from prior work in audio-visual speech enhancement that relies on lip and facial visuals. In this paper, we first demonstrate that egocentric visual information is helpful for noise suppression. We compare object recognition and action classification-based visual feature extractors and investigate methods to align audio and visual representations. Then, we examine different fusion strategies for the aligned features, and locations within the noise suppression model to incorporate visual information. Experiments demonstrate that visual features are most helpful when used to generate additive correction masks. Finally, in order to ensure that the visual features are discriminative with respect to different noise types, we introduce a multi-task learning framework that jointly optimizes audio-visual noise suppression and video-based acoustic event detection. This proposed multi-task framework outperforms the audio-only baseline on all metrics, including a 0.16 PESQ improvement. Extensive ablations reveal the improved performance of the proposed model with multiple active distractors, overall noise types, and across different SNRs. Weipeng He, Ju Lin, Egor Lakomkin, Kaustubh Kalgaonkar |
ICASSP | 3 |
| 2023 | Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding
Suyoun Kim, Akshat Shrivastava, Ju Lin, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 4 |
| 2023 | Directional Speech Recognition for Speaker Disambiguation and Cross-talk Suppression
Ju Lin, Niko Moritz, Ruiming Xie, Kaustubh Kalgaonkar, Christian Fügen, Frank Seide |
INTERSPEECH | 1 |
| 2022 | Architecture for Variable Bitrate Neural Speech Codec with Configurable Computation ComplexityabstractLow bitrate speech codecs have become an area of intense research. Traditional speech codecs, which use signal processing methods to encode and decode speech, often suffer from quality issues at low bitrates. A neural speech codec, which uses a deep neural network in the compression pipeline, can help alleviate this issue. In this paper we present a new neural speech codec that: 1) supports variable bitrates 2) supports packet losses of up to 120 ms and 3) can operate at low-compute and high-compute modes. Our codec uses a hierarchical VQ-VAE (HVQVAE) for encoding and decoding spectral features at different bitrates. The decoded features are fed to a vocoder for speech synthesis. Depending upon the end user’s computing resources, the decoder either uses a powerful WaveRNN or a parametric vocoder for speech synthesis. Our experiments demonstrate that our HVQVAE + WaveRNN setup achieves high audio quality. Tejas Jayashankar, Thilo Köhler, Kaustubh Kalgaonkar, Zhiping Xiu, Jilong Wu, Ju Lin, Prabhav Agrawal |
ICASSP | 6 |
| 2022 | Speech Enhancement for Low Bit Rate Speech CodecabstractSpeech codec compresses the input signal into compact bit stream, which is then decoded at the receiver to generate the best possible perceptual quality. This compression makes storing and transmitting speech efficient. In this work, we propose a neural extension to low bit rate speech codec (e.g., Codec2) that aims to improve the perceptual quality of synthesized speech. Our proposed framework combines decoded audio with neural embeddings without breaking the existing speech coders. In addition to embeddings, we also use the least-square generative adversarial network (LSGAN) to reduce artifacts and prevent over-smoothing in the reconstructed audio. The Mean Opinion Scores (MOS) from the listening tests show that our framework can boost the audio quality of speech encoded at 3.6kbps to outperform that of speech encoded at 6kbps using Opus. Ju Lin, Kaustubh Kalgaonkar |
ICASSP | 1 |
| 2021 | A Time-Domain Convolutional Recurrent Network for Packet Loss ConcealmentabstractPacket loss may affect a wide range of applications that use voice over IP (VoIP), e.g. video conferencing. In this paper, we investigate a time-domain convolutional recurrent network (CRN) for online packet loss concealment. The CRN comprises a convolutional encoder-decoder structure and long short-term memory (LSTM) layers, which have been shown to be suitable for real-time speech enhancement applications. Moreover, we propose lookahead and masked training to further improve the performance of the CRN framework. Experimental results show that the proposed system outperforms a baseline system using only LSTM layers in terms of two objective metrics – perceptual evaluation of speech quality (PESQ) and short-term objective intelligibility (STOI); it also reduces the word error rate (WER) more than the baseline when used as a frontend for speech recognition. The advantage of the proposed system is also verified in a subjective evaluation by the mean opinion score (MOS). Ju Lin, Kaustubh Kalgaonkar, Gil Keren, Didi Zhang, Christian Fügen |
ICASSP | 1 |
| 2021 | Systematic Evaluation and Enhancement of Speech Recognition in Operational Medical EnvironmentsabstractOperational medical environments require reliable hands-free solutions to extract data from audio captured under noisy scenarios during rescue missions and provide timely information. However, approaches using automatic speech recognition (ASR) and natural language processing (NLP) techniques are complex as these conversations have a wide range of noise, involve medical terms from multiple speakers, and occur in high-stress environments, among others. These are further complicated by the lack of large training datasets for operational medical scenarios. To address these issues, we developed a platform that enables resilient hands-free data collection, preserves complete documentation through stages of care, and presents the information in near real-time, critical for the medical operation. Our work uniquely focused on systematic evaluation and improvement of a deep neural network-based ASR system by leveraging realistic testing data obtained from medical simulations of battlefield scenarios, which to our knowledge have not been addressed in any prior work. The system performance is shown to improve significantly using multi-style training, language model adaptation for the medical domain, speech enhancement, and NLP techniques. Snigdhaswin Kar, Prabodh Mishra, Ju Lin, Minjae Woo, Nicholas Deas, Caleb Linduff, Sufeng Niu, Jerome McClendon, D. Hudson Smith, Melissa C. Smith, Ronald W. Gimbel, Kuang-Ching Wang |
IJCNN | 3 |
| 2021 | A Two-Stage Approach to Speech Bandwidth Extension
Ju Lin, Kaustubh Kalgaonkar, Gil Keren, Didi Zhang, Christian Fügen |
Interspeech | 1 |
| 2021 | Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional NetworksabstractMulti-stage learning is an effective technique to invoke multiple deep-learning modules sequentially. This paper applies multi-stage learning to speech enhancement by using a multi-stage structure, where each stage comprises a self-attention (SA) block followed by stacks of temporal convolutional network (TCN) blocks with doubling dilation factors. Each stage generates a prediction that is refined in a subsequent stage. A fusion block is inserted at the input of later stages to re-inject original information. The resulting multi-stage speech enhancement system, in short, multi-stage SA-TCN, is compared with state-of-the-art deep-learning speech enhancement methods using the LibriSpeech and VCTK data sets. The multi-stage SA-TCN system's hyper-parameters are fine-tuned, and the impact of the SA block, the fusion block and the number of stages are determined. The use of a multi-stage SA-TCN system as a front-end for automatic speech recognition systems is investigated as well. It is shown that the multi-stage SA-TCN systems perform well relative to other state-of-the-art systems in terms of speech enhancement and speech recognition scores. Ju Lin, Adriaan J. de Lind van Wijngaarden, Kuang-Ching Wang, Melissa C. Smith |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Improved Speech Enhancement Using a Time-Domain GAN with Mask LearningabstractSpeech enhancement is an essential component in robust automatic speech recognition (ASR) systems. Most speech enhancement methods are nowadays based on neural networks that use feature-mapping or mask-learning. This paper proposes a novel speech enhancement method that integrates time-domain feature mapping and mask learning into a unified framework using a Generative Adversarial Network (GAN). The proposed framework processes the received waveform and decouples speech and noise signals, which are fed into two short-time Fourier transform (STFT) convolution 1-D layers that map the waveforms to spectrograms in the complex domain. These speech and noise spectrograms are then used to compute the speech mask loss. The proposed method is evaluated using the TIMIT data set for seen and unseen signal-to-noise ratio conditions. It is shown that the proposed method outperforms the speech enhancement methods that use Deep Neural Network (DNN) based speech enhancement or a Speech Enhancement Generative Adversarial Network (SEGAN). Ju Lin, Sufeng Niu, Adriaan J. de Lind van Wijngaarden, Jerome McClendon, Melissa C. Smith, Kuang-Ching Wang |
INTERSPEECH | 1 |
| 2019 | Speech Enhancement Using Forked Generative Adversarial Networks with Spectral Subtraction
Ju Lin, Sufeng Niu, Zice Wei, Adriaan J. de Lind van Wijngaarden, Melissa C. Smith, Kuang-Ching Wang |
INTERSPEECH | 1 |
| 2019 | Printed flexible thin-film transistors based on different types of modified liquid metal with good mobility
Ju Lin, Tianying Liu |
Sci. China Inf. Sci. | 2 |
| 2016 | Automatic Pronunciation Evaluation of Non-Native Mandarin Tone by Using Multi-Level Confidence Measures
Ju Lin, Yanlu Xie, Jinsong Zhang 0001 |
INTERSPEECH | 1 |