Xuan Shi

dblp:42/3261 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-1387-5418ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
abstract
We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the RAIL license at https://github.com/tiantiaf0627/voxlect.
Tiantian Feng, Anfeng Xu, Xuan Shi, Thanathai Lertpetchpun, Yoonjeong Lee, Dani Byrd, Shri Narayanan
KDD (1)4
2026 Speech acoustics to rt-MRI articulatory dynamics inversion with video diffusion model
Xuan Shi, Tiantian Feng, Jay Park, Christina Hagedorn, Louis Goldstein, Shri Narayanan
Comput. Speech Lang.1
2025 Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
abstract
Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Resonance Imaging (MRI) videos of the human vocal tract during speech based on arbitrary audio or speech input. We leverage large pre-trained speech models, which are embedded with prior knowledge, to generalize the visual domain to unseen data using an speech-to-video diffusion model. Our findings demonstrate that the visual generation significantly benefits from the pre-trained speech representations. We also observed that evaluating phonemes in isolation is challenging but becomes more straightforward when assessed within the context of spoken words. Limitations of the current results include the presence of unsmooth tongue motion and video distortion when the tongue contacts the palate. The source code is available for the public at: https://github.com/Hong7Cong/SPAN-rtmri.git
Hong Nguyen, Sean Foley, Xuan Shi, Tiantian Feng, Shri Narayanan
ICASSP4
2025 Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
Tiantian Feng, Anfeng Xu, Xuan Shi, Somer Bishop, Shri Narayanan
INTERSPEECH3
2025 Examining Test-Time Adaptation for Personalized Child Speech Recognition
Zhonghao Shi, Xuan Shi, Anfeng Xu, Tiantian Feng, Harshvardhan Srivastava, Shri Narayanan, Maja J. Mataric
INTERSPEECH2
2025 75-Speaker Annot-16: A benchmark dataset for speech articulatory rt-MRI annotation with articulator contours and phonetic alignment
Xuan Shi, Yubin Zhang, Yijing Lu, Marcus Ma, Tiantian Feng, Asterios Toutios, Haley Hsu, Louis Goldstein, Shri Narayanan
INTERSPEECH1
2025 Co-registration of real-time MRI and respiration for speech research
Yubin Zhang, Prakash Kumar, Xuan Shi, Haley Hsu, Shri Narayanan, Krishna S. Nayak, Louis Goldstein
INTERSPEECH5
2024 Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
abstract
Speech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli.We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals.Our approach aims to directly reconstruct listened speech waveforms given EEG signals, where no intermediate acoustic feature processing step is required.The proposed method consists of an EEG module and a speech module along with a connector.The EEG module learns to better represent EEG signals, while the speech module generates speech waveforms from model representations.The connector learns to bridge the distributions of the latent spaces of EEG and speech.The proposed framework is both simple and efficient, by allowing single-step inference, and outperforms prior works on objective metrics.A fine-grained phoneme analysis is conducted to unveil model characteristics of speech decoding.The source code is available here: github.com/lee-jhwn/fesde.
Aditya Kommineni, Tiantian Feng, Kleanthis Avramidis, Xuan Shi, Sudarsana Reddy Kadiri, Shri Narayanan
INTERSPEECH5
2023 Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?
abstract
With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS techniques. In this study, we analyze the shortcomings of a TTS-based MIDI-to-audio system and improve it in terms of feature computation, model selection, and training strategy, aiming to synthesize highly natural-sounding audio. Moreover, we conducted an extensive model evaluation through listening tests, pitch measurement, and spectrogram analysis. This work demonstrates not only synthesis of highly natural music but offers a thorough analytical approach and useful outcomes for the community. Our code, pre-trained models, supplementary materials, and audio samples are open sourced at https://github.com/nii-yamagishilab/midi-to-audio.
Xuan Shi, Erica Cooper, Xin Wang 0037, Junichi Yamagishi, Shri Narayanan
ICASSP1
2022 Use of Speaker Recognition Approaches for Learning and Evaluating Embedding Representations of Musical Instrument Sounds
abstract
Constructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as multi-instrument synthesis and timbre transfer. The framework of Automatic Speaker Verification (ASV) provides us with architectures and evaluation methodologies for verifying the identities of unseen speakers, and these can be repurposed for the task of learning and evaluating a musical instrument sound embedding space that can support unseen instruments. Borrowing from state-of-the-art ASV techniques, we construct a musical instrument recognition model that uses a SincNet front-end, a ResNet architecture, and an angular softmax objective function. Experiments on the NSynth and RWC datasets show our model’s effectiveness in terms of equal error rate (EER) for unseen instruments, and ablation studies show the importance of data augmentation and the angular softmax objective. Experiments also show the benefit of using a CQT-based filterbank for initializing SincNet over a Mel filterbank initialization. Further complementary analysis of the learned embedding space is conducted with t-SNE visualizations and probing classification tasks, which show that including instrument family labels as a multi-task learning target can help to regularize the embedding space and incorporate useful structure, and that meaningful information such as playing style, which was not included during training, is contained in the embeddings of unseen instruments.
Xuan Shi, Erica Cooper, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 A hybrid parallel cellular automata model for urban growth simulation over GPU/CPU heterogeneous architectures
abstract
As an important spatiotemporal simulation approach and an effective tool for developing and examining spatial optimization strategies (e.g., land allocation and planning), geospatial cellular automata (CA) models often require multiple data layers and consist of complicated algorithms in order to deal with the complex dynamic processes of interest and the intricate relationships and interactions between the processes and their driving factors. Also, massive amount of data may be used in CA simulations as high-resolution geospatial and non-spatial data are widely available. Thus, geospatial CA models can be both computationally intensive and data intensive, demanding extensive length of computing time and vast memory space. Based on a hybrid parallelism that combines processes with discrete memory and threads with global memory, we developed a parallel geospatial CA model for urban growth simulation over the heterogeneous computer architecture composed of multiple central processing units (CPUs) and graphics processing units (GPUs). Experiments with the datasets of California showed that the overall computing time for a 50-year simulation dropped from 13,647 seconds on a single CPU to 32 seconds using 64 GPU/CPU nodes. We conclude that the hybrid parallelism of geospatial CA over the emerging heterogeneous computer architectures provides scalable solutions to enabling complex simulations and optimizations with massive amount of data that were previously infeasible, sometimes impossible, using individual computing approaches.
Qingfeng Guan 0001, Xuan Shi, Miaoqing Huang, Chenggang Lai
Int. J. Geogr. Inf. Sci.2