Kevin Tang

dblp:88/9005 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 Automatic Speech Recognition of African American English: Lexical and Contextual Effects
abstract
Automatic Speech Recognition (ASR) models often struggle with the phonetic, phonological, and morphosyntactic features found in African American English (AAE). This study focuses on two key AAE variables: Consonant Cluster Reduction (CCR) and ING-reduction. It examines whether the presence of CCR and ING-reduction increases ASR misrecognition. Subsequently, it investigates whether end-to-end ASR systems without an external Language Model (LM) are more influenced by lexical neighborhood effect and less by contextual predictability compared to systems with an LM. The Corpus of Regional African American Language (CORAAL) was transcribed using wav2vec 2.0 with and without an LM. CCR and ING-reduction were detected using the Montreal Forced Aligner (MFA) with pronunciation expansion. The analysis reveals a small but significant effect of CCR and ING on Word Error Rate (WER) and indicates a stronger presence of lexical neighborhood effect in ASR systems without LMs.
Hamid Mojarad, Kevin Tang
INTERSPEECH2
2025 Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
abstract
Automatic Speech Recognition (ASR) systems struggle with regional dialects due to biased training which favours mainstream varieties. While previous research has identified racial, age, and gender biases in ASR, regional bias remains underexamined. This study investigates ASR performance on Newcastle English, a well-documented regional dialect known to be challenging for ASR. A two-stage analysis was conducted: first, a manual error analysis on a subsample identified key phonological, lexical, and morphosyntactic errors behind ASR misrecognitions; second, a case study focused on the systematic analysis of ASR recognition of the regional pronouns ``yous'' and ``wor''. Results show that ASR errors directly correlate with regional dialectal features, while social factors play a lesser role in ASR mismatches. We advocate for greater dialectal diversity in ASR training data and highlight the value of sociolinguistic analysis in diagnosing and addressing regional biases.
Dana Serditova, Kevin Tang, Jochen Steffens
INTERSPEECH2
2025 Modeling Probabilistic Reduction using Information Theory and Naive Discriminative Learning
abstract
This study compares probabilistic predictors based on information theory with Naive Discriminative Learning (NDL) predictors in modeling acoustic word duration, focusing on probabilistic reduction. We examine three models using the Buckeye corpus: one with NDL-derived predictors using information-theoretic formulas, one with traditional NDL predictors, and one with N-gram probabilistic predictors. Results show that the N-gram model outperforms both NDL models, challenging the assumption that NDL is more effective due to its cognitive motivation. However, incorporating information-theoretic formulas into NDL improves model performance over the traditional model. This research highlights a) the need to incorporate not only frequency and contextual predictability but also average contextual predictability, and b) the importance of combining information-theoretic metrics of predictability and information derived from discriminative learning in modeling acoustic reduction.
Anna Stein, Kevin Tang
INTERSPEECH2
2025 An Efficient Optimization Criterion for Multi-View Feature Representation Learning
abstract
The training of contemporary machine learning (ML) models, particularly deep neural networks (DNNs), often relies on enormous data sources to properly tune model parameters. As a result, achieving competitive results with limited training data and computational resources has been recognized as a significant bottleneck to advance ML. To address these issues, multi-view representation learning has emerged. However, how to efficiently build multi-view learning models remains a big challenge. In this paper, a novel optimization criterion is proposed to tackle this challenge. Specifically, the proposed criterion ensures speedy and effective parameter selection, reducing the effort to reach optimal design of the model while maintaining performance. To validate the efficiency and generalizability of the presented solution, experiments were conducted on face recognition and few-shot learning for image classification using four databases of different scales. Experimental results demonstrate the superiority of the proposed approach, offering an efficient yet robust solution to the data-scarcity challenge.
Lei Gao 0001, Kai Liu 0032, Kevin Tang, Ling Guan
ISM3
2025 Individual differences in language acquisition: The impact of study abroad on native English speakers learning Spanish
abstract
• Studying abroad influences the acquisition of Spanish lenition in native English speakers. • Voicing is the primary factor affecting lenition in both native and non-native speakers. • L2 learners produce fricative-like forms more frequently than approximant-like forms. • Individual differences shape learners' phonological development during immersion. • Studying abroad alone does not guarantee native-like lenition patterns in L2 learners. This study investigated the acquisition of lenition in Spanish voiced stops (/b, d, ɡ/) by native English speakers during a study-abroad program, focusing on individual differences and influencing factors. Lenition, characterized by the weakening of stops into fricative-like ([β], [ð], [ɣ]) or approximant-like ([β̞], [ð̞], [ɣ̞]) forms, poses challenges for L2 learners due to its gradient nature and the absence of analogous approximant forms in English. Results indicated that learners aligned with native speakers in recognizing voicing as the primary cue for lenition, yet their productions diverged, favoring fricative-like over approximant-like realizations. This preference reflects the combined influence of articulatory ease, acoustic salience, and cognitive demands. Individual variability in learners’ trajectories highlights the role of exposure to native input and sociolinguistic engagement. Learners benefitting from richer, informal interactions with native speakers showed greater alignment with native patterns, while others demonstrated more limited progress. However, native input alone was insufficient for learners to internalize subtler distinctions such as place of articulation and stress. These findings emphasize the need for combining immersive experiences with targeted instructional strategies to address articulatory and cognitive challenges. This study contributes to the understanding of L2 phonological acquisition and offers insights for designing more effective language learning programs to support lenition acquisition in Spanish.
Ratree Wayland, Rachel Meyer, Sophia Vellozzi, Kevin Tang
Speech Commun.4
2024 Leveraging Syntactic Dependencies in Disambiguation: The Case of African American English
abstract
African American English (AAE) has received recent attention in the field of natural language processing (NLP). Efforts to address bias against AAE in NLP systems tend to focus on lexical differences. When the unique structures of AAE are considered, the solution is often to remove or neutralize the differences. This work leverages knowledge about the unique linguistic structures to improve automatic disambiguation of habitual and non-habitual meanings of “be” in naturally produced AAE transcribed speech. Both meanings are employed in AAE but examples of Habitual be are rare in already limited AAE data. Generally, representing additional syntactic information improves semantic disambiguation of habituality. Using an ensemble of classical machine learning models with a representation of the unique POS and dependency patterns of Habitual be, we show that integrating syntactic information improves the identification of habitual uses of “be” by about 65 F1 points over a simple baseline model of n-grams, and as much as 74 points. The success of this approach demonstrates the potential impact when we embrace, rather than neutralize, the structural uniqueness of African American English.
Wilermine Previlon, Alice Rozet, Jotsna Gowda, Bill Dyer, Kevin Tang, Sarah Moeller
LREC/COLING5
2024 VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point Correspondence
abstract
Current diffusion-based video editing primarily focuses on structure-preserved editing by utilizing various dense correspondences to ensure temporal consistency and motion alignment. However, these approaches are often in-effective when the target edit involves a shape change. To embark on video editing with shape change, we explore customized video subject swapping in this work, where we aim to replace the main subject in a source video with a target subject having a distinct identity and potentially different shape. In contrast to previous methods that rely on dense correspondences, we introduce the Video Swap framework that exploits semantic point correspondences, inspired by our observation that only a small number of semantic points are necessary to align the subject's motion trajectory and modify its shape. We also introduce various user-point interactions (e.g., removing points and dragging points) to address various semantic point correspondence. Extensive experiments demonstrate state-of-the-art video subject swapping results across a variety of real-world videos.
Yuchao Gu, Yipin Zhou, Bichen Wu, Licheng Yu, Jia-Wei Liu, Rui Zhao 0001, Jay Zhangjie Wu, Junhao Zhang 0001, Zheng Shou 0001, Kevin Tang
CVPR10
2024 Modeling probabilistic reduction across domains with Naive Discriminative Learning
Anna Stein, Kevin Tang
INTERSPEECH2
2023 Computational Storage for an Energy-Efficient Deep Neural Network Training System
Shiju Li 0001, Kevin Tang, Jin Lim, Chul-Ho Lee, Jongryool Kim
Euro-Par2
2023 "It Has to Ignite Their Creativity": Opportunities for Generative Tools for Game Masters
abstract
In this paper, we describe the results of interviews we conducted with experienced game masters (GMs) of tabletop role-playing games. In these interviews they discussed the challenges they face preparing and running game sessions, as well as the tools they use. From these interviews, we used qualitative analysis to discover three guidelines for designing generative tools for GMs: 1. provide inspiration, not answers; 2. allow for customization of the generative possibilities; 3. prioritize ease of use (speed, portability, and online accessibility) in design. While there are some limitations to our approach which we describe in this paper, we found these rules summarized the needs of the GMs we interviewed and provide useful advice for those interested in designing generative tools for GMs.
Kevin Tang, Terra Mae Gasque, Rachel Donley, Anne Sullivan
FDG1
2020 CDeepEx: Contrastive Deep Explanations
Amir Feghahati, Christian R. Shelton, Michael J. Pazzani, Kevin Tang
ECAI4
2020 Gated Story Structure and Dramatic Agency in Sam Barlow's Telling Lies
Terra Mae Gasque, Kevin Tang, Brad Rittenhouse, Janet H. Murray
ICIDS2
2020 Understanding Racial Disparities in Automatic Speech Recognition: The Case of Habitual "be"
Joshua L. Martin, Kevin Tang
INTERSPEECH2
2009 Serving Ads from localhost for Performance, Privacy, and Profit
Saikat Guha 0002, Alexey Reznichenko, Kevin Tang, Hamed Haddadi 0001, Paul Francis
HotNets3