Bowon Lee

dblp:22/3933 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-5417-5699ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Computer networks · 7Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Towards Scalable and Robust Multilingual ASR for Indian Languages with MixLoRA-Whisper
abstract
India exhibits extensive linguistic diversity, with many regional languages and dialects, yet current multilingual automatic speech recognition (ASR) models provide limited support, especially for low-income and rural populations who rely on spoken communication. We apply MixLoRA, a parameterefficient fine-tuning method proposed for large language models, to Whisper to improve ASR performance. MixLoRA employs multiple LoRA experts and dynamically selects the most relevant experts per token, enabling better modeling of linguistic variation. By fine-tuning only up to 25.03 % of the parameters on the RESPIN dataset, which covers eight Indian languages with 33 dialects, it achieves a $4.98 \%$ character error rate (CER) on the read speech, yielding a $7.09 \%$ relative CER reduction over the baseline. Performance improved across all languages in read speech and five in spontaneous speech. These results demonstrate that MixLoRA can effectively enhance ASR for low-resource, dialect-rich languages.
Yeseul Park, Bowon Lee
ASRU2
2025 Improved Recognition of the Speech of People with Parkinson's Who Stutter
abstract
Stuttering is a speech disorder often associated with neurological conditions, including Parkinson’s disease (PD). Despite advancements in modern automatic speech recognition (ASR) technologies, today’s systems still face challenges in accurately recognizing dysarthric speech, particularly when stuttering is present. In this study, we propose a novel stuttered speech data augmentation approach to improve dysarthric speech recognition. We utilize typical speech data from LibriSpeech to generate artificial stuttered speech by applying Voice Activity Detection and Forced Alignment techniques to accurately identify word boundaries, and integrating an adaptive stuttering filter to simulate severe stuttering patterns. Additionally, dysarthric speech data from individuals with PD, collected by the Speech Accessibility Project (SAP), is integrated into the model. Our experimental results demonstrate that the proposed augmentation approach outperforms existing methods in enhancing the recognition of stuttered speech. Furthermore, fine-tuning the ASR systems with SAP data yields additional performance improvements for both stuttering and non-stuttering individuals with PD.
Jonghwan Na, Xiuwen Zheng 0003, Bowon Lee, Mark Hasegawa-Johnson
ICASSP3
2025 Cohort-Sensitive Labeling: An Effective Approach for Enhancing ASR Performance
abstract
This paper proposes a cohort-sensitive labeling (CSL) for automatic speech recognition (ASR). CSL is a method that distinguishes data labels based on cohorts, allowing models to learn cohort-specific information. For evaluation, we applied CSL using gender information in the training data of LibriSpeech dataset. Experimental results demonstrate that the CSL-based approach outperforms methods without CSL, given sufficient training data. Specifically, our method achieved average word error rate reduction (WERR) of 1.81% on the LibriSpeech test-clean and 5.76% on test-other datasets, when more than 100 hours of data were used for training. Moreover, on TIMIT and Common Voice test sets, it achieved WERR of up to 11.52% and 2.91%, respectively demonstrating its robustness and generalizability to unseen data. Additionally, the proposed method reached up to 97.21% accuracy in classifying the gender cohort, suggesting that ASR models trained with the CSL effectively leverage the cohort information.
Jonghwan Na, Mark Hasegawa-Johnson, Bowon Lee
ICASSP3
2025 Fine-tuning Strategies for Automatic Speech Recognition of Low-Resource Speech with Autism Spectrum Disorder
Yeseul Park, Bowon Lee
INTERSPEECH2
2024 Beyond superficial emotion recognition: Modality-adaptive emotion recognition system
Dohee Kang, Dae Ha Kim, Taein Kim, Bowon Lee, Deok-Hwan Kim, Byung Cheol Song
Expert Syst. Appl.5
2022 Integration of Pre-Trained Networks with Continuous Token Interface for End-to-End Spoken Language Understanding
abstract
Most End-to-End (E2E) Spoken Language Understanding (SLU) networks leverage the pre-trained Automatic Speech Recognition (ASR) networks but still lack the capability to understand the semantics of utterances, crucial for the SLU task. To solve this, recently proposed studies use pre-trained Natural Language Understanding (NLU) networks. However, it is not trivial to fully utilize both pre-trained networks; many solutions were proposed, such as Knowledge Distillation (KD), cross-modal shared embedding and network integration with Interface. We propose a simple and robust integration method for the E2E SLU network with a novel Interface, Continuous Token Interface (CTI). CTI is a junctional representation of the ASR and NLU networks when both networks are pre-trained with the same vocabulary. Thus, we can train our SLU network in an E2E manner without additional modules, such as Gumbel-Softmax. We evaluate our model using SLURP, a challenging SLU dataset and achieve state-of-the-art scores on intent classification and slot filling tasks. We also verify that the NLU network, pre-trained with Masked Language Model (MLM), can utilize a noisy textual representation of CTI. Moreover, we train our model with extra data, SLURP-Synth, and get better results.
Seunghyun Seo, Donghyun Kwak, Bowon Lee
ICASSP3
2022 Style Transfer Using Optimal Transport Via Wasserstein Distance
abstract
Universal style transfer has been proven to be effective through CNN models and VGG networks. However, how well to apply the algorithm’s style is a separate issue. This problem is especially evident in high-resolution images in which case the division and color at the boundary lines are more complex than low-resolution images. For the WCT model, if the image segment is smaller than the filter size, it will be blurred. High-resolution images have much larger number of small segments and the existing WCT model cannot render them clearly. This paper proposes two methods. It uses the Wasserstein distance-based optimal transport so that the resulting style distribution is the same up to the secondary statistics when a style image is applied to the content image, and proposes collaborative distillation, a method to overcome the encoder and decoder dependency of the WCT module. We propose a module that combines these two methods to apply subtle style transfer even to high-resolution images.
Oseok Ryu, Bowon Lee
ICIP2
2022 Phase Vocoder For Time Stretch Based On Center Frequency Estimation
Bowon Lee
INTERSPEECH2
2020 Multi-Attention Multimodal Sentiment Analysis
abstract
Sentiment analysis plays an important role in natural-language processing. It has been performed on multimodal data including text, audio, and video. Previously conducted research does not make full utilization of such heterogeneous data. In this study, we propose a model of Multi-Attention Recurrent Neural Network (MA-RNN) for performing sentiment analysis on multimodal data. The proposed network consists of two attention layers and a Bidirectional Gated Recurrent Neural Network (BiGRU). The first attention layer is used for data fusion and dimensionality reduction, and the second attention layer is used for the augmentation of BiGRU to capture key parts of the contextual information among utterances. Experiments on multimodal sentiment analysis indicate that our proposed model achieves the state-of-the-art performance of 84.31% accuracy on the Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis (CMU-MOSI) dataset. Furthermore, an ablation study is conducted to evaluate the contributions of different components of the network. We believe that our findings of this study may also offer helpful insights into the design of models using multimodal data.
Bowon Lee
ICMR2
2019 Real-Time Object Identification with a Smartphone Knock
abstract
We propose Knocker, a real-time object identification technique with smartphones. Knocker leverages unique impulse signals that are generated by knocking on an object with a smartphone. Knocker does not require any special augmentation for both smartphones and objects.
Taesik Gong, Hyunsung Cho, Bowon Lee, Sung-Ju Lee 0001
MobiSys3
2017 A Survey of Sound Source Localization Methods in Wireless Acoustic Sensor Networks
abstract
Wireless acoustic sensor networks (WASNs) are formed by a distributed group of acoustic-sensing devices featuring audio playing and recording capabilities. Current mobile computing platforms offer great possibilities for the design of audio-related applications involving acoustic-sensing nodes. In this context, acoustic source localization is one of the application domains that have attracted the most attention of the research community along the last decades. In general terms, the localization of acoustic sources can be achieved by studying energy and temporal and/or directional features from the incoming sound at different microphones and using a suitable model that relates those features with the spatial location of the source (or sources) of interest. This paper reviews common approaches for source localization in WASNs that are focused on different types of acoustic features, namely, the energy of the incoming signals, their time of arrival (TOA) or time difference of arrival (TDOA), the direction of arrival (DOA), and the steered response power (SRP) resulting from combining multiple microphone signals. Additionally, we discuss methods not only aimed at localizing acoustic sources but also designed to locate the nodes themselves in the network. Finally, we discuss current challenges and frontiers in this field.
Maximo Cobos, Fabio Antonacci, Anastasios Alexandridis, Athanasios Mouchtaris, Bowon Lee
Wirel. Commun. Mob. Comput.5
2017 Wireless Acoustic Sensor Networks and Applications
Maximo Cobos, Fabio Antonacci, Athanasios Mouchtaris, Bowon Lee
Wirel. Commun. Mob. Comput.4
2017 Acoustic Sensor Self-Localization: Models and Recent Results
abstract
The wide availability of mobile devices with embedded microphones opens up opportunities for new applications based on acoustic sensor localization (ASL). Among them, this paper highlights mobile device self-localization relying exclusively on acoustic signals, but with previous knowledge of reference signals and source positions. The problem of finding the sensor position is stated as a function of estimated times-of-flight (TOFs) or time-differences-of-flight (TDOFs) from the sound sources to the target microphone, and the main practical issues involved in TOF estimation are discussed. Least-squares ASL solutions are introduced, followed by other strategies inspired by sound source localization solutions: steered-response power, which improves localization accuracy, and a new region-based search, which alleviates complexity. A set of complementary techniques for further improvement of TOF/TDOF estimates are reviewed: sliding windows, matching pursuit, and TOF selection. The paper proceeds with proposing a novel ASL method that combines most of the previous material, whose performance is assessed in a real-world example: in a typical lecture room, the method achieves accuracy better than 20 cm.
Diego B. Haddad, Markus V. S. Lima, Wallace A. Martins, Luiz W. P. Biscainho, Leonardo O. Nunes, Bowon Lee
Wirel. Commun. Mob. Comput.6
2016 Robust Acoustic Self-Localization of Mobile Devices
abstract
Self-localization of smart portable devices serves as foundation for several novel applications. This work proposes a set of algorithms that enable a mobile device to passively determine its position relative to a known reference with centimeter precision, based exclusively on the capture of acoustic signals emitted by controlled sources around it. The proposed techniques tackle typical practical issues such as reverberation, unknown speed of sound, line-of-sight obstruction, clock skew, and the need for asynchronous operation. After their theoretical developments and off-line simulations, the methods are assessed as real-time applications embedded into off-the-shelf mobile devices operating in real scenarios. When line of sight is available, position estimation errors are at most 4 cm using recorded signals.
Diego B. Haddad, Wallace A. Martins, Maurício do Vale Madeira da Costa, Luiz W. P. Biscainho, Leonardo O. Nunes, Bowon Lee
IEEE Trans. Mob. Comput.6
2015 A Volumetric SRP with Refinement Step for Sound Source Localization
abstract
This letter proposes an efficient method based on the steered-response power (SRP) technique for sound source localization using microphone arrays: the refined volumetric SRP (RV-SRP). By deploying a sparser volumetric grid, the RV-SRP achieves a significant reduction of the computational complexity without sacrificing the accuracy of location estimates. In addition, a refinement step improves on the compromise between complexity and accuracy. Experiments conducted in both simulated- and real-data scenarios show that the RV-SRP outperforms state-of-the-art methods in accuracy with lower computational cost.
Markus V. S. Lima, Wallace A. Martins, Leonardo O. Nunes, Luiz W. P. Biscainho, Tadeu N. Ferreira, Maurício do Vale Madeira da Costa, Bowon Lee
IEEE Signal Process. Lett.7
2015 Multimodal Multi-Channel On-Line Speaker Diarization Using Sensor Fusion Through SVM
abstract
Speaker diarization (SD) is the process of assigning speech segments of an audio stream to its corresponding speakers, thus comprising the problem of voice activity detection (VAD), speaker labeling/identification, and often sound source localization (SSL). Most research activities in the past aimed towards applications as broadcast news, meetings, conversational telephony, and automatic multimodal data annotation, where SD may be performed off-line. However, a recent research focus is human-computer interaction (HCI) systems where SD must be performed on-line, and in real-time, as in modern gaming devices and interaction with large displays. Often, such applications further suffer from noise, reverberations, and overlapping speech, making them increasingly challenging. In such situations, multimodal/multisensory approaches can provide more accurate results than unimodal ones, given a data stream may compensate for occasional instabilities of other modalities. Accordingly, this paper presents an on-line multimodal SD algorithm designed to work in a realistic environment with multiple, overlapping speakers. Our work employs a microphone array, a color camera, and a depth sensor as input streams, from which speech-related features are extracted to be later merged through a support vector machine approach consisting of VAD and SSL modules. Speaker identification is incorporated through a hybrid technique of face positioning history and face recognition. Our final SD approach experimentally achieves an average diarization error rate of 11.48% in scenarios with up to three simultaneous speakers, and is able to run 3.2 × real-time.
Vicente P. Minotto, Cláudio R. Jung, Bowon Lee
IEEE Trans. Multim.3
2014 Multiview image and video interpolation using weighted vector median filters
abstract
In Depth Image-Based Rendering (DIBR), interpolated views generated using one or two cameras usually present artifacts and holes due to occlusions and/or inconsistencies in the input disparity maps. In this paper we propose a multiple (3 or more) camera view interpolation technique that is able to combine redundant projections in a single interpolated view by using Weighted Vector Median Filters (WVMFs). By expressing the weights of the WVMF using both the distance from each reference view to the synthetic view and a measure of consonance of each projection to the others, we achieve a high quality view interpolation without holes and common visual artifacts, such as cracks and ghost effects. Additionally, we present an extension to multiview video sequences by imposing temporal coherence in the estimated disparity maps.
Guilherme P. Fickel, Cláudio R. Jung, Bowon Lee
ICIP3
2014 Simultaneous-Speaker Voice Activity Detection and Localization Using Mid-Fusion of SVM and HMMs
abstract
Humans can extract speech signals that they need to understand from a mixture of background noise, interfering sound sources, and reverberation for effective communication. Voice Activity Detection (VAD) and Sound Source Localization (SSL) are the key signal processing components that humans perform by processing sound signals received at both ears, sometimes with the help of visual cues by locating and observing the lip movements of the speaker. Both VAD and SSL serve as the crucial design elements for building applications involving human speech. For example, systems with microphone arrays can benefit from these for robust speech capture in video conferencing applications, or for speaker identification and speech recognition in Human Computer Interfaces (HCIs). The design and implementation of robust VAD and SSL algorithms in practical acoustic environments are still challenging problems, particularly when multiple simultaneous speakers exist in the same audiovisual scene. In this work we propose a multimodal approach that uses Support Vector Machines (SVMs) and Hidden Markov Models (HMMs) for assessing the video and audio modalities through an RGB camera and a microphone array. By analyzing the individual speakers' spatio-temporal activities and mouth movements, we propose a mid-fusion approach to perform both VAD and SSL for multiple active and inactive speakers. We tested the proposed algorithm in scenarios with up to three simultaneous speakers, showing an average VAD accuracy of 95.06% with an average error of 10.9 cm when estimating the three-dimensional locations of the speakers.
Vicente P. Minotto, Cláudio R. Jung, Bowon Lee
IEEE Trans. Multim.3
2013 Determining co-location using a sequential hypothesis test on patterns of silence
abstract
In everyday meetings, automatic association of co-located mobile devices would ease sharing of web-links, media, and other information. We propose a method that compares patterns of silence from device microphones to detect co-location of those devices. This method works with unsynchronized audio capture, requires only 100bps and preserves privacy. We show how to formulate pattern matching in a sequential hypothesis framework so that changes in co-location status (when people leave or join a meeting) can be determined promptly, and how to compute the likelihood ratio in practice. Using 16 hours of captured audio, we show that our approach can correctly determine device co-location with a low error rate of 0.05%, and can detect co-location changes 10 seconds faster than a similar decision rule based on a constant time window. Compared to a prior audio signature method, we achieve higher accuracy at 1/7 the bit rate.
Wai-tian Tan, Ramin Samadani, Bowon Lee, Mary Baker
ICASSP3
2013 Sensing device co-location through patterns of silence
abstract
This document describes the technology behind the accompanying video, which gives a demonstration of determining the dynamic group membership of a meeting by matching patterns of relative audio silence, or "silence signatures," sensed by mobile devices.
Wai-tian Tan, Mary Baker, Bowon Lee, Ramin Samadani
MobiSys3
2013 The sound of silence
abstract
A list of the dynamically changing group membership of a meeting supports a variety of meeting-related activities. Effortless content sharing might be the most important application, but we can also use it to provide business card information for attendees, feed information into calendar applications to simplify scheduling of follow-up meetings, populate the membership of collaborative editing applications, mailing lists, and social networks, and perform many other tasks.
Wai-tian Tan, Mary Baker, Bowon Lee, Ramin Samadani
SenSys3
2012 Voice activity detection and speaker localization using audiovisual cues
Dante A. Blauth, Vicente P. Minotto, Cláudio R. Jung, Bowon Lee, Ton Kalker
Pattern Recognit. Lett.4
2012 On the quality-assessment of reverberated speech
Amaro A. de Lima, Thiago de M. Prego, Sergio L. Netto, Bowon Lee, Amir Said, Ronald W. Schafer, Ton Kalker, Majid Fozunbal
Speech Commun.4
2012 A Parametric Objective Quality Assessment Tool for Speech Signals Degraded by Acoustic Echo
abstract
This paper discusses the automatic quality assessment of echo-degraded speech in the context of teleconference systems. Subjective listening tests conducted over a carefully designed database of signals degraded by acoustic echo have been used to assess how this impairment is perceived and to determine which parameters have a significant impact on speech quality. The results have shown that, similarly to electric transmission line echo, acoustic echo is mainly influenced by echo delay and echo gain. Based on this observation, a mapping between these two parameters and the mean subjective score is devised. Moreover, a signal-based algorithm for the estimation of these parameters is described, and its performance is evaluated. The complete system comprising both the parameter estimators and the mapping function achieves a correlation of 94% between predicted and actual subjective scores, and can be employed as a non-intrusive monitoring tool for in-service quality evaluation of teleconference systems. Further validation indicates the operating range of the proposed quality assessment tool can be extended by proper retraining.
Leonardo O. Nunes, Flávio R. Avila, Alan Freihof Tygel, Luiz W. P. Biscainho, Bowon Lee, Amir Said, Ronald W. Schafer
IEEE Trans. Speech Audio Process.5
2011 Correlogram template matching for time-delay estimation
abstract
We propose a correlogram-based time delay estimation method using signals modeled as the output of the cochlea, where the low-level signal processing happens in the human auditory system. With a normalized correlogram that preserves time-delay patterns that are invariant to speech features such as formants, we employ two-dimensional template matching for time-delay estimation. Experimental results show that our method outperforms a traditional correlogram-based method as well as the GCC-PHAT, especially for short analysis windows in a moderately reverberant environment.
Bowon Lee, Ton Kalker, Ronald W. Schafer
ICASSP1
2011 A system approach to residual echo suppression in robust hands-free teleconferencing
abstract
This paper presents a system approach to the residual echo suppression (RES) problem in a noisy acoustic environment. We propose a method that takes advantage of our existing robust acoustic echo cancellation system in order to obtain a residual echo estimate that closely resembles the true, noise-free residual echo. To achieve improved RES during strong near-end interference (e.g., double talk), a psychoacoustic postfilter is also used. The simulation results show that our RES based on the system approach outperforms a conventional estimation method. Comparing the postfiltered output to the unprocessed one indicates that our proposed RES approach can raise the PESQ score by more than half a point.
Jason Wung, Ted S. Wada, Biing-Hwang Juang, Bowon Lee, Ton Kalker, Ronald W. Schafer
ICASSP4
2011 Robust user context analysis for multimodal interfaces
abstract
Multimodal Interfaces that enable natural means of interaction using multiple modalities such as touch, hand gestures, speech, and facial expressions represent a paradigm shift in human-computer interfaces. Their aim is to allow rich and intuitive multimodal interaction similar to human-to-human communication and interaction. From the multimodal system's perspective, apart from the various input modalities themselves, user context information such as states of attention and activity, and identities of interacting users can help greatly in improving the interaction experience. For example, when sensors such as cameras (webcams, depth sensors etc.) and microphones are always on and continuously capturing signals in their environment, user context information is very useful to distinguish genuine system-directed activity from ambient speech and gesture activity in the surroundings, and distinguish the "active user" from among a set of users. Information about user identity may be used to personalize the system's interface and behavior -- e.g. the look of the GUI, modality recognition profiles, and information layout -- to suit the specific user. In this paper, we present a set of algorithms and an architecture that performs audiovisual analysis of user context using sensors such as cameras and microphone arrays, and integrates components for lip activity and audio direction detection (speech activity), face detection and tracking (attention), and face recognition (identity). The proposed architecture allows the component data flows to be managed and fused with low latency, low memory footprint, and low CPU load, since such a system is typically required to run continuously in the background and report events of attention, activity, and identity, in real-time, to consuming applications.
Muthuselvam Selvaraj, Bowon Lee
ICMI3
2011 Degradation Type Classifier for Full Band Speech Contaminated With Echo, Broadband Noise, and Reverberation
abstract
This paper addresses the problem of identifying impairment types that might be present in a speech signal. In particular, three acoustically induced degradation types that occur in teleconference systems are considered: acoustic echo, reverberation, and broadband noise, as well as combinations among them. The proposed system is double-ended (full reference) and is developed using a database of degraded full-band speech signals created according to a model for teleconference systems. A set of features obtained from both the degraded and non-degraded signals is proposed and shown to adequately capture information associated with each degradation type. A random forest classifier and a support vector machine are successfully employed, achieving a classification error below 2%. Such classifiers can be used to select an appropriate quality assessment tool for a given degraded signal.
Leonardo O. Nunes, Luiz W. P. Biscainho, Bowon Lee, Amir Said, Ton Kalker, Ronald W. Schafer
IEEE ACM Trans. Audio Speech Lang. Process.3
2010 Spectral entropy-based voice activity detector for videoconferencing systems
Bowon Lee, Debargha Muhkerjee
INTERSPEECH1
2010 Massively parallel processing of signals in dense microphone arrays
abstract
Arrays with large number of microphones can be very effective on audio processing tasks, like denoising, acoustic echo removal, etc. New microphone technologies enable creating large arrays with very low cost per component, but the system can still be very expensive due to costs of transmitting all signals to a single processor, and the computational resources to process the large amount of data. We show how a massively-parallel signal processing approach can solve the cost issues, when applied to the problem of sound source localization. We consider the case where each microphone is coupled to simple processing circuitry, which have full-bandwidth access to data from a few other microphones, while only shared power and low-bandwidth connections are provided between each microphone and a central processor. We discuss implementation issues, and show experimental results obtained in simulations and in microphone array measurements.
Amir Said, Ton Kalker, Bowon Lee, Majid Fozunbal
ISCAS3
2009 An objective method for quality assessment of ultra-wideband speech corrupted by echo
abstract
Modern telepresence systems can deliver multimedia signals of unprecedentedly high quality of experience to the user. Setting and maintaining such services call for reliable and automatic tools for multimedia quality probing, in special those targeted at speech data along the transmission path. Most of the objective methods for sound quality assessment (QA) in the literature are intended for either speech signals of 4- to 8-kHz bandwidth or general audio until 24 kHz, but are not specifically designed for speech at high sampling-rates. This work approaches quality evaluation of full-band (24 kHz) high-quality speech corrupted by echo. A simple metric singled out from a standardized double-ended tool for audio QA is proposed as a solution for the problem at hand. Quality measures from a set of speech stimuli corrupted by echo under controlled conditions were obtained via listening tests to allow calibration and evaluation of the proposed method. Experimental results reveal an overall correlation of 0.94 between objective and subjective scores, even in the presence of moderate additive noise.
Luiz W. P. Biscainho, Paulo Antonio Andrade Esquef, Fabio P. Freeland, Leonardo O. Nunes, Alan Freihof Tygel, Bowon Lee, Amir Said, Ton Kalker, Ronald W. Schafer
MMSP6
2009 Quality assessment of audio: Increasing applicability scope of objective methods via prior identification of impairment types
abstract
In this paper the design of a double-ended (intrusive) diagnostic tool for identifying five types of degradation in audio signals is reported. The impairment types taken into consideration are additive contamination with pink noise, occurrences of signal mutes, distortion by magnitude clipping, and the previous two types mixed with pink noise. As a simple solution to accomplish the established goal, a threshold-based hierarchical classification system is proposed, being completely defined from pre-processing of the input signals, passing through the estimation of a few characteristic features, up to data clustering criteria. Performance evaluation of the classifier is carried out via a validation database containing 60 impaired signals for each type of impairment, with five distinct degradation intensity levels. Considering the types and range of degradation levels considered in this work, excellent results are achieved, scoring above 96% of correctly classified data in the worst case. System performance in identifying mixed impairment types tends to deteriorate as the strength of the noise component increases.
Paulo Antonio Andrade Esquef, Luiz W. P. Biscainho, Leonardo O. Nunes, Bowon Lee, Amir Said, Ton Kalker, Ronald W. Schafer
MMSP4
2009 Feature analysis for quality assessment of reverberated speech
abstract
This paper analyzes the ability of several measurements to quantify the reverberation effect in speech signals. We consider an intrusive scheme, in which the clean and reverberated signals are available, allowing one to estimate the corresponding room impulse response (RIR) signal. An artificial neural network (ANN) is trained for all features and used in a regression approach to estimate the human perceptual evaluation in a mean opinion score (MOS) 1–5 scale. Dimensionality reduction approaches are applied to generate a simpler ANN regression, establishing the most representative features for the problem at hand. A correlation level of 85% with subjective test scores was achieved by reducing the input-vector dimension from 10 to 3, including only the features of reverberation time, room spectral variance, and direct-to-reverberant energy ratio.
Amaro A. de Lima, Thiago de M. Prego, Sergio L. Netto, Bowon Lee, Amir Said, Ronald W. Schafer, Ton Kalker, Majid Fozunbal
MMSP4
2009 ConnectBoard: A remote collaboration system that supports gaze-aware interaction and sharing
abstract
We present ConnectBoard, a new system for remote collaboration where users experience natural interaction with one another, seemingly separated only by a vertical, transparent sheet of glass. It overcomes two key shortcomings of conventional video communication systems: the inability to seamlessly capture natural user interactions, like using hands to point and gesture at parts of shared documents, and the inability of users to look into the camera lens without taking their eyes off the display. We solve these problems by placing the camera behind the screen, where the remote user is virtually located. The camera sees through the display to capture images of the user. As a result, our setup captures natural, frontal views of users as they point and gesture at shared media displayed on the screen between them. Users also never have to take their eyes off their screens to look into the camera lens. Our novel optical solution based on wavelength multiplexing can be easily built with off-the-shelf components and does not require custom electronics for projector-camera synchronization.
Kar-Han Tan, Ian N. Robinson, Ramin Samadani, Bowon Lee, Dan Gelb, Alex Vorbau, W. Bruce Culbertson, John G. Apostolopoulos
MMSP4
2008 On the quality assessment of sound signals
abstract
This paper constitutes an introduction to the field of quality evaluation of sound (speech and audio) signals. The need for such an assessment is inherent to modern communications: VoIP, mobile phone, or teleconference systems require meaningful measures of performance, which may ultimately assure good service or profitable business. A brief survey on subjective and objective evaluation methods is provided. Recent developments as well as new topics to be investigated are also addressed. Experiments are conducted to illustrate how to validate quality assessment methods.
Amaro A. de Lima, Fabio P. Freeland, Rafael A. de Jesus, Bruno C. Bispo, Luiz W. P. Biscainho, Sergio L. Netto, Amir Said, Ton Kalker, Ronald W. Schafer, Bowon Lee, Mehrban Jam
ISCAS10
2004 AVICAR: audio-visual speech corpus in a car environment
abstract
We describe a large audio-visual speech corpus recorded in a car environment, as well as the equipment and procedures used to build this corpus. Data are collected through a multi-sensory array consisting of eight microphones on the sun visor and four video cameras on the dashboard. The script for the corpus consists of four categories: isolated digits, isolated letters, phone numbers, and sentences, all in English. Speakers from various language backgrounds are included, 50 male and 50 female. In order to vary the signal-to-noise ratio, each script has five different noise conditions: idling, driving at 35 mph with windows open and closed, and driving at 55 mph with windows open and closed. The corpus is available through
Bowon Lee, Mark Hasegawa-Johnson, Camille Goudeseune, Suketu Kamdar, Sarah Borys, Ming Liu 0009, Thomas S. Huang
INTERSPEECH1