Suzy J. Styles

dblp:239/0549 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
12since 2021 · last 2024
0000-0003-3517-9680ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2024 A preregistered investigation of language-specific distributional learning advantages in English-Mandarin bilingual adults
Hannah L. Goh, Suzy J. Styles
CogSci2
2024 Missing /y/: Vowel perception in bilinguals whose languages differ in whether the high front rounded vowel is phonemic
Suzy J. Styles
CogSci2
2023 Do you say what you hear? Perception-production link of a phoneme contrast in Singapore Mandarin
Hannah L. Goh, Suzy J. Styles
CogSci2
2023 Addressing the Word Gap in Singapore: Growth Trajectory of Vocabulary of Children Growing up in Bi/Multilingual Households
Fei Ting Woon, Suzy J. Styles
CogSci2
2023 MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarization
abstract
To enhance the reliability and robustness of language identification (LID) and language diarization (LD) systems for heterogeneous populations and scenarios, there is a need for speech processing models to be trained on datasets that feature diverse language registers and speech patterns. We present the MERLIon CCS challenge, featuring a first-of-its-kind Zoom video call dataset of parent-child shared book reading, of over 30 hours with over 300 recordings, annotated by multilingual transcribers using a high-fidelity linguistic transcription protocol. The audio corpus features spontaneous and in-the-wild English-Mandarin code-switching, child-directed speech in non-standard accents with diverse language-mixing patterns recorded in a variety of home environments. This report describes the corpus, as well as LID and LD results for our baseline and several systems submitted to the MERLIon CCS challenge using the corpus.
Yi Han Victoria Chua, Hexin Liu, L. Paola García-Perera, Fei Ting Woon, Jinyi Wong, Xiangyu Zhang 0005, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels, Suzy J. Styles
INTERSPEECH10
2023 Investigating model performance in language identification: beyond simple error statistics
abstract
Language development experts need tools that can automatically identify languages from fluent, conversational speech and provide reliable estimates of usage rates at the level of an individual recording. However, LID systems are typically evaluated on metrics such as equal error rate and balanced accuracy, applied at the level of an entire speech corpus. These overview metrics do not provide information about model performance at the level of individual speakers, recordings, or units of speech with different linguistic characteristics. Overview statistics may mask systematic errors in model performance for some subsets of the data, and consequently, have worse performance on data derived from some subsets of human speakers, creating a kind of algorithmic bias. Here, we investigate how well a number of LID systems perform on individual recordings and speech units with different linguistic properties in the MERLIon CCS Challenge featuring accented code-switched child-directed speech.
Suzy J. Styles, Yi Han Victoria Chua, Fei Ting Woon, Hexin Liu, L. Paola García-Perera, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels
INTERSPEECH1
2022 Perception of a phoneme contrast in Singaporean English-Mandarin bilingual adults: A preregistered study of individual differences
Hannah L. Goh, Suzy J. Styles
CogSci2
2022 'Sheep' and 'Ship': An investigation into English vowel merger in multilingual Singapore
Fei Ting Woon, Suzy J. Styles
CogSci2
2022 PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language Identification
abstract
We propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training.In this model, named PHO-LID, a self-supervised phoneme segmentation task and a LID task share a convolutional neural network (CNN) module, which encodes both language identity and sequential phonemic information in the input speech to generate an intermediate sequence of "phonotactic" embeddings.These embeddings are then fed into transformer encoder layers for utterance-level LID.We call this architecture CNN-Trans.We evaluate it on AP17-OLR data and the MLS14 set of NIST LRE 2017, and show that the PHO-LID model with multitask optimization exhibits the highest LID performance among all models, achieving over 40% relative improvement in terms of average cost on AP17-OLR data compared to a CNN-Trans model optimized only for LID.The visualized confusion matrices imply that our proposed method achieves higher performance on languages of the same cluster in NIST LRE 2017 data than the CNN-Trans model.A comparison between predicted phoneme boundaries and corresponding audio spectrograms illustrates the leveraging of phoneme information for LID.
Hexin Liu, L. Paola García-Perera, Andy W. H. Khong, Suzy J. Styles, Sanjeev Khudanpur
INTERSPEECH4
2021 A preregistered study exploring language-specific distributional learning advantages in English-Mandarin bilingual adults
Hannah L. Goh, Suzy J. Styles
CogSci2
2021 From Alien Zoo to Spy School: A Preregistered Study of Linguistic Sound Symbolism and its Links to Reading in 8-year-olds
Fei Ting Woon, De-Fu Yap, Shirong Cai, Evelyn C. Law, Lourdes Mary Daniel, Suzy J. Styles
CogSci6
2021 End-to-End Language Diarization for Bilingual Code-Switching Speech
abstract
We propose two end-to-end neural configurations for language diarization on bilingual code-switching speech. The first, a BLSTM-E2E architecture, includes a set of stacked bidirectional LSTMs to compute embeddings and incorporates the deep clustering loss to enforce grouping of languages belonging to the same class. The second, an XSA-E2E architecture, is based on an x-vector model followed by a self-attention encoder. The former encodes frame-level features into segmentlevel embeddings while the latter considers all those embeddings to generate a sequence of segment-level language labels. We evaluated the proposed methods on the dataset obtained from the shared task B in WSTCSMC 2020 and our handcrafted simulated data from the SEAME dataset. Experimental results show that our proposed XSA-E2E architecture achieved a relative improvement of 12.1% in equal error rate and a 7.4% relative improvement on accuracy compared with the baseline algorithm in the WSTCSMC 2020 dataset. Our proposed XSA-E2E architecture achieved an accuracy of 89.84% with a baseline of 85.60% on the simulated data derived from the SEAME dataset.
Hexin Liu, L. Paola García-Perera, Justin Dauwels, Andy W. H. Khong, Sanjeev Khudanpur, Suzy J. Styles
Interspeech7
2018 Pre-Readers at the Alien Zoo: A Preregistered Study of the Predictors of Dyslexia and Linguistic Sound Symbolism in 6-year-olds
Fei Ting Woon, Yap-Seng Chong, Lourdes Mary Daniel, Birit F. P. Broekman, Shirong Cai, Suzy J. Styles
CogSci6