VLDB 2026 Research / reviewers in the wild / expert
Shuju Shi
dblp:153/0774
· DBLP profile ↗
10ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-8349-1145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | L2 Vowel Acquisition Analysis at the Inventory LevelabstractTheories for second language acquisition (SLA) of phonology/phonetics and pronunciation/accent often resort to the similarity/dissimilarity between sound inventories of the first language (L1) and the second language (L2). In this study, we investigate the acquisition of English vowels at the inventory level by learners with six different L1 backgrounds, each with a distinct relationship to English. We use Pillai scores as quantitative measures of categorization and generalized additive mixed models (GAMMs) to analyze time-dynamic formant contours of diphthongs. Our approach emphasizes the importance of studying language acquisition at the inventory level and highlights the need for analyzing both the phonological representation and phonetic realization of learners’ L1 phone inventories. Shuju Shi |
ASRU | 1 |
| 2025 | What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech DetectionabstractRecent advances in text-to-speech technology have enabled highly realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation.While audio antispoofing systems are critical for detecting such threats, prior research has predominantly focused on acoustic-level perturbations, leaving the impact of linguistic variation largely unexplored.In this paper, we investigate the linguistic sensitivity of both open-source and commercial anti-spoofing detectors by introducing TAPAS (Transcript-to-Audio Perturbation Anti-Spoofing), a novel framework for transcript-level adversarial attacks.Our extensive evaluation shows that even minor linguistic perturbations can significantly degrade detection accuracy: attack success rates exceed 60% on several open-source detector-voice pairs, and the accuracy of one commercial detector drops from 100% on synthetic audio to just 32%.Through a comprehensive feature attribution analysis, we find that linguistic complexity and model-level audio embedding similarity are key factors contributing to detector vulnerabilities.To illustrate the real-world risks, we replicate a recent Brad Pitt audio deepfake scam and demonstrate that TAPAS can bypass commercial detectors.These findings underscore the need to move beyond purely acoustic defenses and incorporate linguistic variation into the design of robust anti-spoofing systems.Our source code is available at https: //github.com/nqbinh17/audio_linguist ic_adversarial. Shuju Shi, Ryan Ofman, Thai Le |
EMNLP | 2 |
| 2025 | Neutral Tone Variation in Beijing Mandarin: Is Neutral Tone Toneless?
Fengming Liu, Chien-Jer Charles Lin, Monica Nesbitt, Shuju Shi |
INTERSPEECH | 5 |
| 2023 | Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation ScoringabstractRecent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation quality representation. In this paper, we propose to use linguistic-acoustic similarity to explicitly measure the deviation of non-native production from its native reference for pronunciation assessment. Specifically, the deviation is first estimated by the cosine similarity between reference phone embedding and corresponding acoustic embedding. Next, a phone-level Goodness of pronunciation (GOP) pre-training stage is introduced to guide this similarity-based learning for better initialization of the aforementioned two embeddings. Finally, a transformer-based hierarchical pronunciation scorer is used to map a sequence of phone embeddings, acoustic embeddings along with their similarity measures to predict the final utterance-level score. Experimental results on the non-native databases suggest that the proposed system significantly outperforms the baselines, where the acoustic and phone embeddings are simply added or concatenated. A further examination shows that the phone embeddings learned in the proposed approach are able to capture linguistic-acoustic attributes of native pronunciation as references. Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee |
ICASSP | 4 |
| 2023 | An ASR-Free Fluency Scoring Approach with Self-Supervised LearningabstractA typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for the subsequent calculation of fluency-related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free approach for automatic fluency assessment using self-supervised learning (SSL). Specifically, wav2vec2.0 is used to extract frame-level speech features, followed by K-means clustering to assign a pseudo label (cluster index) to each frame. A BLSTM-based model is trained to predict an utterance-level fluency score from frame-level SSL features and the corresponding cluster indexes. Neither speech transcription nor time stamp information is required in the proposed system. It is ASR-free and can potentially avoid the ASR errors effect in practice. Experimental results carried out on non-native English databases show that the proposed approach significantly improves the performance in the "open response" scenario as compared to previous methods and matches the recently reported performance in the "read aloud" scenario. Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee |
ICASSP | 4 |
| 2023 | Phonetic and Prosody-aware Self-supervised Learning Approach for Non-native Fluency Scoring
Kaiqi Fu, Shaojun Gao, Shuju Shi, Xiaohai Tian, Wei Li 0119, Zejun Ma 0001 |
INTERSPEECH | 3 |
| 2023 | Disentangling the Contribution of Non-native Speech in Automated Pronunciation Assessment
Shuju Shi, Kaiqi Fu, Yiwei Gu, Xiaohai Tian, Shaojun Gao, Wei Li 0119, Zejun Ma 0001 |
INTERSPEECH | 1 |
| 2019 | Capturing L1 Influence on L2 Pronunciation by Simulating Perceptual Space Using Acoustic Features
Shuju Shi, Chilin Shih, Jinsong Zhang 0001 |
INTERSPEECH | 1 |
| 2016 | Automatic Assessment and Error Detection of Shadowing Speech: Case of English Spoken by Japanese Learners
Shuju Shi, Yosuke Kashiwagi, Shohei Toyama, Junwei Yue, Yutaka Yamauchi, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2014 | A preliminary study on acoustic correlates of tone2+tone2 disyllabic word stress in Mandarin
Shuju Shi, Jinsong Zhang 0001 |
INTERSPEECH | 2 |