Jun Wang 0037

dblp:125/8189-37 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0001-7265-217XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3
YearPublicationVenuePosition
2025 Articulatory Vowel Distinctiveness in Spanish
Kristin Teplansky, Emily Rangel, Mimi LaValley, Beiming Cao, Jun Wang 0037
INTERSPEECH6
2024 Direct Speech Synthesis from Non-Invasive, Neuromagnetic Signals
David F. Harwath, Debadatta Dash, Paul Ferrari, Jun Wang 0037
INTERSPEECH5
2022 Data Augmentation for End-to-end Silent Speech Recognition for Laryngectomees
Beiming Cao, Kristin Teplansky, Nordine Sebkhi, Arpan Bhavsar, Omer T. Inan, Robin Samlan, Ted Mau, Jun Wang 0037
INTERSPEECH8
2021 Investigating Speech Reconstruction for Laryngectomees for Silent Speech Interfaces
Beiming Cao, Nordine Sebkhi, Arpan Bhavsar, Omer T. Inan, Robin Samlan, Ted Mau, Jun Wang 0037
Interspeech7
2021 Investigating the Utility of Multimodal Conversational Technology and Audiovisual Analytic Measures for the Assessment and Monitoring of Amyotrophic Lateral Sclerosis at Scale
abstract
We propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of the various acoustic and visual speech metrics automatically extracted by the platform. 82 healthy controls and 54 people with ALS (pALS) were instructed to interact with the platform and completed a battery of speaking tasks designed to probe the acoustic, articulatory, phonatory, and respiratory aspects of their speech. We find that multiple acoustic (rate, duration, voicing) and visual (higher order statistics of the jaw and lip) speech metrics show statistically significant differences between controls, bulbar symptomatic and bulbar pre-symptomatic patients. We report on the sensitivity and specificity of these metrics using five-fold cross-validation. We further conducted a LASSO-LARS regression analysis to uncover the relative contributions of various acoustic and visual features in predicting the severity of patients' ALS (as measured by their self-reported ALSFRS-R scores). Our results provide encouraging evidence of the utility of automatically extracted audiovisual analytics for scalable remote patient assessment and monitoring in ALS.
Michael Neumann 0001, Oliver Roesler, Jackson Liscombe, Hardik Kothare, David Suendermann-Oeft, David Pautler, Indu Navar, Aria Anvar, Jochen Kumm, Raquel Norel, Ernest Fraenkel, Alexander V. Sherman, James D. Berry, Gary L. Pattee, Jun Wang 0037, Jordan R. Green, Vikram Ramanarayanan
Interspeech15
2020 Decoding Speech Evoked Jaw Motion from Non-invasive Neuromagnetic Oscillations
abstract
Speech decoding-based brain-computer interfaces (BCIs) are the next-generation neuroprostheses that have the potential for real-time communication assistance to patients with locked-in syndrome (fully paralyzed but aware). Recent invasive speech decoding studies have demonstrated the possibility of speech kinematics decoding, where articulatory movements were decoded from the brain activity signals for speech synthesis, as an alternative solution to direct brain-to-speech mapping. As a starting point toward a non-invasive speech-neuroprosthesis, in this study, we investigated the decoding of continuous jaw kinematic trajectories directly from non-invasive neuromagnetic signals during speech production. The compensatory jaw behavior exhibited by patients with amyotrophic lateral sclerosis (ALS) is prevalent, hence, accurate decoding of the jaw kinematics could be a path for developing efficient communicative BCIs for these patients. Using magnetoencephalography (MEG), we recorded brain signals and jaw motions simultaneously from four subjects as they spoke short phrases. We trained a long short-term memory (LSTM) regression model to successfully map the brain activity to jaw motion with about 0.80 average correlation score across all four subjects. In addition, we also examined the decoding performance of specific frequency bands within the neural signals and found that the Delta (0.3-4Hz) and high-gamma (62-125 Hz and 125-250 Hz) frequencies independently can account for the major contributions in jaw motion decoding. Experimental results indicated that the jaw kinematics can be successfully decoded from non-invasive neural (MEG) signals.
Debadatta Dash, Paul Ferrari, Jun Wang 0037
IJCNN3
2020 Neural Speech Decoding for Amyotrophic Lateral Sclerosis
Debadatta Dash, Paul Ferrari, Angel W. Hernandez-Mulero, Daragh Heitzman, Sara G. Austin, Jun Wang 0037
INTERSPEECH6
2020 Increasing the Intelligibility and Naturalness of Alaryngeal Speech Using Voice Conversion and Synthetic Fundamental Frequency
Tuan Dinh, Alexander Kain, Robin Samlan, Beiming Cao, Jun Wang 0037
INTERSPEECH5
2020 Tongue and Lip Motion Patterns in Alaryngeal Speech
Kristin Teplansky, Alan Wisler, Beiming Cao, Wendy Liang, Chad W. Whited, Ted Mau, Jun Wang 0037
INTERSPEECH7
2019 Spatial and Spectral Fingerprint in the Brain: Speaker Identification from Single Trial MEG Signals
Debadatta Dash, Paul Ferrari, Jun Wang 0037
INTERSPEECH3
2019 Towards a Speaker Independent Speech-BCI Using Speaker Adaptation
Debadatta Dash, Alan Wisler, Paul Ferrari, Jun Wang 0037
INTERSPEECH4
2019 Early Identification of Speech Changes Due to Amyotrophic Lateral Sclerosis Using Machine Classification
Sarah E. Gutz, Jun Wang 0037, Yana Yunusova, Jordan R. Green
INTERSPEECH2
2018 Automatic Early Detection of Amyotrophic Lateral Sclerosis from Intelligible Speech Using Convolutional Neural Networks
Kwanghoon An, Myung Jong Kim, Kristin Teplansky, Jordan R. Green, Thomas F. Campbell, Yana Yunusova, Daragh Heitzman, Jun Wang 0037
INTERSPEECH8
2018 Articulation-to-Speech Synthesis Using Articulatory Flesh Point Sensors' Orientation Information
Beiming Cao, Myung Jong Kim, Jun R. Wang, Jan P. H. van Santen, Ted Mau, Jun Wang 0037
INTERSPEECH6
2018 Automatic Speech Recognition with Articulatory Information and a Unified Dictionary for Hindi, Marathi, Bengali and Oriya
Debadatta Dash, Myung Jong Kim, Kristin Teplansky, Jun Wang 0037
INTERSPEECH4
2018 Dysarthric Speech Recognition Using Convolutional LSTM Neural Network
Myung Jong Kim, Beiming Cao, Kwanghoon An, Jun Wang 0037
INTERSPEECH4
2017 Towards decoding speech production from single-trial magnetoencephalography (MEG) signals
abstract
Patients with locked-in-syndrome (fully paralyzed but aware) struggle in their life and communication. Providing a level of communication offers these patients a chance to resume a meaningful life. Current brain-computer interface (BCI) communication requires users to build words from single letters selected on a screen, which is extremely inefficient. Faster approaches for their speech communication are highly needed. This project investigated the possibility to decode spoken phrases from non-invasive brain activity (MEG) signals. This direct brain-to-text mapping approach may provide a significantly faster communication rate than current BCIs can provide. We used dynamic time warping and Wiener filtering for noise reduction and then Gaussian mixture model and artificial neural network as the decoders. Preliminary results showed the possibility of decoding speech production from non-invasive brain signals. The best phrase classification accuracy was up to 94.54% from single-trial whole-head MEG recordings.
Jun Wang 0037, Myung Jong Kim, Angel W. Hernandez-Mulero, Daragh Heitzman, Paul Ferrari
ICASSP1
2017 Integrating Articulatory Information in Deep Learning-Based Text-to-Speech Synthesis
Beiming Cao, Myung Jong Kim, Jan P. H. van Santen, Ted Mau, Jun Wang 0037
INTERSPEECH5
2017 Multiview Representation Learning via Deep CCA for Silent Speech Recognition
Myung Jong Kim, Beiming Cao, Ted Mau, Jun Wang 0037
INTERSPEECH4
2017 Generalizing DTW to the multi-dimensional case requires an adaptive approach
Mohammad Shokoohi-Yekta, Bing Hu 0001, Hongxia Jin, Jun Wang 0037, Eamonn J. Keogh
Data Min. Knowl. Discov.4
2017 Speaker-Independent Silent Speech Recognition From Flesh-Point Articulatory Movements Using an LSTM Neural Network
abstract
Silent speech recognition (SSR) converts non-audio information such as articulatory movements into text. SSR has the potential to enable persons with laryngectomy to communicate through natural spoken expression. Current SSR systems have largely relied on speaker-dependent recognition models. The high degree of variability in articulatory patterns across different speakers has been a barrier for developing effective speaker-independent SSR approaches. Speaker-independent SSR approaches, however, are critical for reducing the amount of training data required from each speaker. In this paper, we investigate speaker-independent SSR from the movements of flesh points on tongue and lip with articulatory normalization methods that reduce the inter-speaker variation. To minimize the across-speaker physiological differences of the articulators, we propose Procrustes matching-based articulatory normalization by removing locational, rotational, and scaling differences. To further normalize the articulatory data, we apply feature-space maximum likelihood linear regression and i-vector. In this paper, we adopt a bidirectional long short term memory recurrent neural network (BLSTM) as an articulatory model to effectively model the articulatory movements with long-range articulatory history. A silent speech data set with flesh points was collected using an electromagnetic articulograph (EMA) from twelve healthy and two laryngectomized English speakers. Experimental results showed the effectiveness of our speaker-independent SSR approaches on healthy as well as laryngectomy speakers. In addition, BLSTM outperformed standard deep neural network. The best performance was obtained by BLSTM with all the three normalization approaches combined.
Myung Jong Kim, Beiming Cao, Ted Mau, Jun Wang 0037
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Dysarthric Speech Recognition Using Kullback-Leibler Divergence-Based Hidden Markov Model
Myung Jong Kim, Jun Wang 0037, Hoirin Kim
INTERSPEECH2
2016 Towards Automatic Detection of Amyotrophic Lateral Sclerosis from Speech Acoustic and Articulatory Samples
Jun Wang 0037, Prasanna V. Kothalkar, Beiming Cao, Daragh Heitzman
INTERSPEECH1
2015 Parkinson's condition estimation using speech acoustic and inversely mapped articulatory data
Seongjun Hahm, Jun Wang 0037
INTERSPEECH2
2015 Speaker-independent silent speech recognition with across-speaker articulatory normalization and speaker adaptive training
Jun Wang 0037, Seongjun Hahm
INTERSPEECH1
2015 Accelerating Dynamic Time Warping Clustering with a Novel Admissible Pruning Strategy
abstract
Clustering time series is a useful operation in its own right, and an important subroutine in many higher-level data mining analyses, including data editing for classifiers, summarization, and outlier detection. While it has been noted that the general superiority of Dynamic Time Warping (DTW) over Euclidean Distance for similarity search diminishes as we consider ever larger datasets, as we shall show, the same is not true for clustering. Thus, clustering time series under DTW remains a computationally challenging task. In this work, we address this lethargy in two ways. We propose a novel pruning strategy that exploits both upper and lower bounds to prune off a large fraction of the expensive distance calculations. This pruning strategy is admissible; giving us provably identical results to the brute force algorithm, but is at least an order of magnitude faster. For datasets where even this level of speedup is inadequate, we show that we can use a simple heuristic to order the unavoidable calculations in a most-useful-first ordering, thus casting the clustering as an anytime algorithm. We demonstrate the utility of our ideas with both single and multidimensional case studies in the domains of astronomy, speech physiology, medicine and entomology.
Nurjahan Begum, Liudmila Ulanova, Jun Wang 0037, Eamonn J. Keogh
KDD3
2015 On the Non-Trivial Generalization of Dynamic Time Warping to the Multi-Dimensional Case
abstract
In the last decade, Dynamic Time Warping (DTW) has emerged as the distance measure of choice for virtually all time series data mining applications. This is the result of significant progress in improving DTW's efficiency, and multiple empirical studies showing that DTW-based classifiers at least equal the accuracy of all their rivals across dozens of datasets. Thus far, most of the research has considered only the one-dimensional case, with practitioners generalizing to the multi-dimensional case in one of two ways. In general, it appears the community believes either that the two ways are equivalent, or that the choice is irrelevant. In this work, we show that this is not the case. The two most commonly used multidimensional DTW methods can produce different classifications, and neither one dominates over the other. This seems to suggest that one should learn the best method for a particular application. However, we will show that this is not necessary; a simple, principled rule can be used on a case-by-case basis to predict which of the two methods we should give credence to. Our method allows us to ensure that classification results are at least as accurate as the better of the two rival methods, and in many cases, our method is strictly more accurate. We demonstrate our ideas with the most extensive set of multi-dimensional time series classification experiments ever attempted.
Mohammad Shokoohi-Yekta, Jun Wang 0037, Eamonn J. Keogh
SDM2
2014 Opti-speech: a real-time, 3d visual feedback system for speech training
abstract
We describe an interactive 3D system to provide talkers with real-time information concerning their tongue and jaw movements during speech. Speech movement is tracked by a magnetometer system (Wave; NDI, Waterloo, Ontario, Canada). A customized interface allows users to view their current tongue position (represented as an avatar consisting of flesh-point markers and a modeled surface) placed in a synchronously moving, transparent head. Subjects receive augmented visual feedback when tongue sensors achieve the correct place of articulation. Preliminary data obtained for a group of adult talkers suggest this system can be used to reliably provide real-time feedback for American English consonant place of articulation targets. Future studies, including tests with communication disordered subjects, are described.
William F. Katz, Thomas F. Campbell, Jun Wang 0037, Eric Farrar, Jessie Colette Eubanks, Arvind Balasubramanian, B. Prabhakaran 0001, Rob Rennaker
INTERSPEECH3
2014 Contribution of tongue lateral to consonant production
abstract
Speaking requires the coordinated movements of individual articulators. Understanding each articulator’s contribution to speech is fundamental not only for understanding how speech is produced, but also for optimizing speech assessment and treatment. Our recent work has studied the individual contributions of tongue tip, tongue blade, tongue body front, tongue body back, upper lip, and lower lip movement to speech sound production by tracking the motion of sensors attached on the midline of tongue and lips. An optimal set of articulators (tongue tip, tongue body back, upper lip, and lower lip) has been found. However, the tongue lateral (side)’s contribution to speech is still poorly understood. We therefore investigated the contribution of the tongue lateral region to consonant production by analyzing the motion of a sensor attached to the side of tongue. Repeated productions of 12 consonants (including the lateral approximant /l/) were collected from six native English speakers. Consonant classification accuracy based on articulatory movement data obtained using a support vector machine was used as an indication of contribution level. The results suggest that sagittal movement of the tongue lateral sensor did not significantly benefit consonant classification, over and above the optimal set. Implications of these findings are discussed.
Jun Wang 0037, William F. Katz, Thomas F. Campbell
INTERSPEECH1
2014 Across-speaker articulatory normalization for speaker-independent silent speech recognition
abstract
Silent speech interfaces (SSIs), which recognize speech from articulatory information (i.e., without using audio information), have the potential to enable persons with laryngectomy or a neurological disease to produce synthesized speech with a natural sounding voice using their tongue and lips. Current approaches to SSIs have largely relied on speaker-dependent recognition models to minimize the negative effects of talker variation on recognition accuracy. Speaker-independent approaches are needed to reduce the large amount of training data required from each user; only limited articulatory samples are often available for persons with moderate to severe speech impairments, due to the logistic difficulty of data collection. This paper reported an across-speaker articulatory normalization approach based on Procrustes matching, a bidimensional regression technique for removing translational, scaling, and rotational effects of spatial data. A dataset of short functional sentences was collected from seven English talkers. A support vector machine was then trained to classify sentences based on normalized tongue and lip movements. Speaker-independent classification accuracy (tested using leave-one-subject-out cross validation) improved significantly, from 68.63 % to 95.90%, following normalization. These results support the feasibility of a speaker-independent SSI using Procrustes matching as the basis for articulatory normalization across speakers. Index Terms: silent speech recognition, speech kinematics, Procrustes analysis, support vector machine
Jun Wang 0037, Ashok Samal, Jordan R. Green
INTERSPEECH1
2013 Individual articulator's contribution to phoneme production
abstract
Speech sounds are the result of coordinated movements of individual articulators. Understanding each articulator's role in speech is fundamental not only for understanding how speech is produced, but also for optimizing speech assessments and treatments. In this paper, we studied the individual contributions of six articulators, tongue tip, tongue blade, tongue body front, tongue body back, upper lip, and lower lip to phoneme classification. A total of 3,838 vowel and consonant production samples were collected from eleven native English speakers. The results of speech movement classification using a support vector machine indicated that the tongue encoded significantly more information than lips, and that the tongue tip may be the most important single articulator among all of the six for phoneme production. Furthermore, our results suggested that the tracking of four articulators (i.e., tongue tip, tongue body back, upper lip, and lower lip) may be sufficient for distinguishing major English phonemes based on articulatory movements.
Jun Wang 0037, Jordan R. Green, Ashok Samal
ICASSP1
2013 SMASH: a tool for articulatory data processing and analysis
abstract
Recent innovations in 3D motion capture technology such as electromagnetic articulography (EMA) are providing unprecedented access to the intricate movements of the articulators during speech production. Although these technological advances afford exciting opportunities for advancing the assessment and treatment of speech, they have presented new challenges associated with data collection, processing, and analysis. To address these challenges, we have standardized our EMA data collection protocols and developed a Matlab-based software tool, SMASH, for processing, visualizing, and analyzing speech movement data. The goal of the software is to advance research on speech production by improving the efficiency and reliability of speech movement analyses.
Jordan R. Green, Jun Wang 0037, David L. Wilson
INTERSPEECH2
2012 Sentence recognition from articulatory movements for silent speech interfaces
abstract
Recent research has demonstrated the potential of using an articulation-based silent speech interface for command-and-control systems. Such an interface converts articulation to words that can then drive a text-to-speech synthesizer. In this paper, we have proposed a novel near-time algorithm to recognize whole-sentences from continuous tongue and lip movements. Our goal is to assist persons who are aphonic or have a severe motor speech impairment to produce functional speech using their tongue and lips. Our algorithm was tested using a functional sentence data set collected from ten speakers (3012 utterances). The average accuracy was 94.89% with an average latency of 3.11 seconds for each sentence prediction. The results indicate the effectiveness of our approach and its potential for building a real-time articulation-based silent speech interface for clinical applications.
Jun Wang 0037, Ashok Samal, Jordan R. Green, Frank Rudzicz
ICASSP1
2012 Whole-Word Recognition from Articulatory Movements for Silent Speech Interfaces
abstract
Articulation-based silent speech interfaces convert silently produced speech movements into audible words. These systems are still in their experimental stages, but have significant potential for facilitating oral communication in persons with laryngectomy or speech impairments. In this paper, we report the result of a novel, real-time algorithm that recognizes whole-words based on articulatory movements. This approach differs from prior work that has focused primarily on phoneme-level recognition based on articulatory features. On average, our algorithm missed 1.93 words in a sequence of twenty-five words with an average latency of 0.79 seconds for each word prediction using a data set of 5,500 isolated word samples collected from ten speakers. The results demonstrate the effectiveness of our approach and its potential for building a real-time articulation-based silent speech interface for health applications.
Jun Wang 0037, Ashok Samal, Jordan R. Green, Frank Rudzicz
INTERSPEECH1
2011 Quantifying Articulatory Distinctiveness of Vowels
abstract
The articulatory distinctiveness among vowels has been frequently characterized descriptively based on tongue height and front-back position; however, very few empirical methods have been proposed to characterize vowels based on time-varying articulatory characteristics. Such information is not only needed to improve knowledge about the articulation of vowels but also to determine the contribution of articulatory imprecision to poor speech intelligibility. In this paper, a novel statistical shape analysis was used to derive a vowel space that depicted the quantified articulatory distinctiveness among vowels based on tongue and lip movements. The effectiveness of the approach was supported by vowel classification accuracy of up to 91.7%. The theoretical relevance and clinical implication of the derived vowel space were discussed. Index Terms: speech production, articulatory vowel space, Procrustes analysis, multi-dimensional scaling
Jun Wang 0037, Jordan R. Green, Ashok Samal, David Marx
INTERSPEECH1