VLDB 2026 Research / reviewers in the wild / expert
Dongmei Wang
dblp:65/4883
· DBLP profile ↗
57ranked-venue papers
26as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 13 first-author · 12 since 2021Artificial intelligence and machine learning · 20 · 11 first-author · 12 since 2021Computer networks · 13 · 4 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A slope-sequence-based deep learning framework for springback prediction and compensation in multi-point stretch forming
Dongmei Wang, Renwei Wang, Changliang Zhang, Yueteng Zhou |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Optimal Topology Control and Energy-Efficient Scheduling Strategy for UAV-Assisted Data Collection in Wireless Sensor NetworksabstractIn recent years, the integration of Unmanned Aerial Vehicles (UAV) into Wireless Sensor Network (WSNs) architectures has emerged as a promising solution for enhancing data collection efficiency. UAV-assisted data collection for WSNs should resolve the problem of energy-efficient communication between UAV and arbitrarily dispersed ground sensors. To enhance the reliability and efficiency of the entire system, an Optimal Topology Control and Energy-efficient Scheduling strategy (OTCES) for UAV-assisted data collection in WSNs is put forward. Based on the clustered-WSNs, a multi-objective model for data collection that minimizes the energy consumption of UAV and sensor nodes is established. Firstly, a density based spatial clustering of applications with noise algorithm is employed to group sensor nodes within the UAV’s effective communication range. Subsequently, to reduce the energy consumption of sensors and data transmission during UAV hovering, an improved deep deterministic policy gradient algorithm is proposed to address the issue of optimal UAV hovering positions and sensor’s transmission power. Then, to minimize the flight energy consumption of UAV, an ant colony optimization-inspired algorithm is designed to figure out the best traverse path. The results of simulation experiments show that compared with traditional methods, the proposed strategy has achieved remarkable results in terms of system energy consumption. Xiaoxue Feng, Dongmei Wang, Junyong Li |
IEEE Internet Things J. | 2 |
| 2025 | Multimodal Deep Learning for Retinal Disease Diagnosis
Dongmei Wang, Yuansong Cai, Wanli Qiao, Menglei Liu |
ISNN | 1 |
| 2025 | CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow MatchingabstractGenerating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model overlapping speech, and synthesize coherent conversations efficiently. In this paper, we introduce CoVoMix2, a fully non-autoregressive framework for zero-shot multi-talker dialogue generation. CoVoMix2 directly predicts mel-spectrograms from multi-stream transcriptions using a flow-matching-based generative model, eliminating the reliance on intermediate token representations. To better capture realistic conversational dynamics, we propose transcription-level speaker disentanglement, sentence-level alignment, and prompt-level random masking strategies. Our approach achieves state-of-the-art performance, outperforming strong baselines like MoonCast and Sesame in speech quality, speaker consistency, and inference speed. Notably, CoVoMix2 operates without requiring transcriptions for the prompt and supports controllable dialogue generation, including overlapping speech and precise timing control, demonstrating strong generalizability to real-world speech generation scenarios. Audio samples are available at https://www.microsoft.com/en-us/research/project/covomix/covomix2. Leying Zhang, Yao Qian, Xiaofei Wang 0009, Manthan Thakker, Dongmei Wang, Jianwei Yu 0001, Yuxuan Hu 0003, Jinyu Li 0001, Yanmin Qian, Sheng Zhao 0002 |
NeurIPS | 5 |
| 2025 | NK-DCHS: An adaptive hybrid immune model for imbalanced anomaly detection
Dongmei Wang, Jinan Gu, Chengwang Xie |
Expert Syst. Appl. | 2 |
| 2025 | Order Matters: The Effect of "AR-First" or "AR-Later" on Consumer Decision-MakingabstractAR-based product displays are widely used across various applications. While there is extensive research on the comparative advantages of AR over traditional displays, a critical gap persists in understanding the effects of the AR order (AR-first or AR-later) on consumer decision-making. This study reveals that, compared to AR-later display, AR-first display significantly reduces decision-making difficulty. This effect is mediated by cognitive load, moderated by purchase motive, AR vividness and need for cognition. For products associated with hedonic motives, AR with high vividness, and consumers with low need for cognition, the impact of AR order on consumer decision-making difficulty is stronger. Consequently, AR-first can yield many benefits, such as reducing cognitive load and decision-making difficulty. These findings align with and enrich the SEAD framework in AR marketing, offering valuable theoretical and practical insights for further exploring the dynamic mechanisms of AR marketing and optimization of practical applications. Chunhua Sun, Dongmei Wang, Ye-Zheng Liu 0001 |
Int. J. Hum. Comput. Interact. | 2 |
| 2025 | SGO: An innovative oversampling approach for imbalanced datasets using SVM and genetic algorithms
Dongmei Wang, Jinan Gu |
Inf. Sci. | 2 |
| 2024 | Profile-Error-Tolerant Target-Speaker Voice Activity DetectionabstractTarget-Speaker Voice Activity Detection (TS-VAD) utilizes a set of speaker profiles alongside an input audio signal to perform speaker diarization. While its superiority over conventional methods has been demonstrated, the method can suffer from errors in speaker profiles, as those profiles are typically obtained by running a traditional clustering-based diarization method over the input signal. This paper proposes an extension to TS-VAD, called Profile-Error-Tolerant TS- VAD (PETTSVAD), which is robust to such speaker profile errors. This is achieved by employing transformer-based TS-VAD that can handle a variable number of speakers and further introducing a set of additional pseudo-speaker profiles to handle speakers undetected during the first pass diarization. During training, we use speaker profiles estimated by multiple different clustering algorithms to reduce the mismatch between the training and testing conditions regarding speaker profiles. Experimental results show that PET-TSVAD consistently outperforms the existing TS-VAD method on both the VoxConverse and DIHARD-I datasets. Dongmei Wang, Naoyuki Kanda, Midia Yousefi, Takuya Yoshioka |
ICASSP | 1 |
| 2024 | TransVIP: Speech to Speech Translation System with Voice and Isochrony PreservationabstractThere is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition, machine translation and text-to-speech models. The primary challenges stem from the inherent complexities involved in direct translation tasks and the scarcity of data. In this study, we introduce a novel model framework TransVIP that leverages diverse datasets in a cascade fashion yet facilitates end-to-end inference through joint probability. Furthermore, we propose two separated encoders to preserve the speaker’s voice characteristics and isochrony from the source speech during the translation process, making it highly suitable for scenarios such as video dubbing. Our experiments on the French-English language pair demonstrate that our model outperforms the current state-of-the-art speech-to-speech translation model. Chenyang Le, Yao Qian, Dongmei Wang, Shujie Liu 0001, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Sheng Zhao 0002, Michael Zeng 0001 |
NeurIPS | 3 |
| 2024 | CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker ConversationsabstractRecent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. CoVoMix first converts dialogue text into multiple streams of discrete tokens, with each token stream representing semantic information for individual talkers. These token streams are then fed into a flow-matching based acoustic model to generate mixed mel-spectrograms. Finally, the speech waveforms are produced using a HiFi-GAN model. Furthermore, we devise a comprehensive set of metrics for measuring the effectiveness of dialogue modeling and generation. Our experimental results show that CoVoMix can generate dialogues that are not only human-like in their naturalness and coherence but also involve multiple talkers engaging in multiple rounds of conversation. This is exemplified by instances generated in a single channel where one speaker's utterance is seamlessly mixed with another's interjections or laughter, indicating the latter's role as an attentive listener. Audio samples are enclosed in the supplementary. Leying Zhang, Yao Qian, Shujie Liu 0001, Dongmei Wang, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Lei He 0005, Sheng Zhao 0002, Michael Zeng 0001 |
NeurIPS | 5 |
| 2024 | Investigating Neural Audio Codecs For Speech Language Model-Based Speech GenerationabstractNeural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the codec system affects the speech generation performance of the SLM. In this work, we examine codec tokens within SLM framework for speech generation to provide insights for effective codec design. We retrain existing high-performing neural codec models on the same data set and loss functions to compare their performance in a uniform setting. We integrate codec tokens into two SLM systems: masked-based parallel speech generation system and an auto-regressive (AR) plus non-auto-regressive (NAR) model-based system. Our findings indicate that better speech reconstruction in codec systems does not guarantee improved speech generation in SLM. A high-quality codec decoder is crucial for natural speech production in SLM, while speech intelligibility depends more on quantization mechanism. Jiaqi Li 0030, Dongmei Wang, Xiaofei Wang 0009, Yao Qian, Shujie Liu 0001, Midia Yousefi, Canrun Li, Chung-Hsien Tsai, Jun-Kun Chen, Sheng Zhao 0002, Jinyu Li 0001, Zhizheng Wu 0001, Michael Zeng 0001 |
SLT | 2 |
| 2024 | KEFSAR: A Solar-Aware Routing Strategy For Rechargeable IoT Based On High-Accuracy PredictionabstractAbstract The high energy density of solar energy gives wireless sensor networks advantages in outdoor monitoring applications. However, long-term stable monitoring is challenging due to frequent weather changes, shading by buildings and trees, etc. The existing research usually uses two technologies to solve the above problems: (1) the energy prediction algorithm, and (2) the energy-aware routing strategy. However, in an actual deployment, frequent weather changes can significantly reduce the accuracy of the existing prediction algorithms. When using the algorithms as the support for energy-aware routing, the network lifetime is less than ideal. The existing routing strategies are in need of further improvement. Because of its lack of environmental adaptability, nodes consume energy quickly and have a high mortality rate. Therefore, aiming at the long-term stability of solar wireless sensor networks, this paper proposes a prediction algorithm based on classification and recurrent neural networks, and integrates the shadow judgement method from our previous research to correct the predicted values. Furthermore, we propose a routing optimization model that can flexibly adjust the target according to the solar intensity. The experimental results show that the prediction and routing scheduling algorithm can significantly improve the energy prediction accuracy (30–50%) and prolong the network lifetime (10–42%) in outdoor small sensor scenarios. Dongchao Ma, Dongmei Wang, Xiaofu Huang, Yuekun Hu, Li Ma 0007 |
Comput. J. | 2 |
| 2024 | Natural gas pipeline leak diagnosis based on manifold learning
Yunqiu Fu, Zhongrui Hu, Dongmei Wang |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Fs-yolo: fire-smoke detection based on improved YOLOv7
Dongmei Wang, Zhongrui Hu, Yongkang Chai |
Multim. Syst. | 1 |
| 2024 | Ea-yolo: efficient extraction and aggregation mechanism of YOLO for fire detection
Dongmei Wang, Dandi Yang, Tianhong yan |
Multim. Syst. | 1 |
| 2023 | Target Sound Extraction with Variable Cross-Modality CluesabstractAutomatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources. It often uses a model conditioned on a fixed form of target sound clues, such as a sound class label, which limits the ways in which users can interact with the model to specify the target sounds. To leverage variable number of clues cross modalities available in the inference phase, including a video, a sound event class, and a text caption, we propose a unified transformer-based TSE model architecture, where a multi-clue attention module integrates all the clues across the modalities. Since there is no off-the-shelf benchmark to evaluate our proposed approach, we build a dataset1based on public corpora, Audioset and AudioCaps. Experimental results for seen and unseen target-sound evaluation sets show that our proposed TSE model can effectively deal with a varying number of clues which improves the TSE performance and robustness against partially compromised clues. Chenda Li, Yao Qian, Zhuo Chen 0006, Dongmei Wang, Takuya Yoshioka, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001 |
ICASSP | 4 |
| 2023 | Target Speaker Voice Activity Detection with Transformers and Its Integration with End-To-End Neural DiarizationabstractThis paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model’s drawback of being unable to handle an arbitrary number of speakers, we investigate model architectures that use input tensors with variable-length time and speaker dimensions. Transformer layers are applied to the speaker axis to make the model output insensitive to the order of the speaker profiles provided to the TS-VAD model. Time-wise sequential layers are interspersed between these speaker-wise transformer layers to allow the temporal and cross-speaker correlations of the input speech signal to be captured. We also extend a diarization model based on end-to-end neural diarization with encoder-decoder based attractors (EEND-EDA) by replacing its dot-product-based speaker detection layer with the transformer-based TS-VAD. Experimental results on VoxConverse show that using the transformers for the cross-speaker modeling reduces the diarization error rate (DER) of TS-VAD by 11.3%, achieving a new state-of-the-art (SOTA) DER of 4.57%. Also, our extended EEND-EDA reduces DER by 6.9% on the CALLHOME dataset relative to the original EEND-EDA with a similar model size, achieving a new SOTA DER of 11.18% under a widely used training data setting. Dongmei Wang, Naoyuki Kanda, Takuya Yoshioka, Jian Wu 0027 |
ICASSP | 1 |
| 2023 | Adapting Multi-Lingual ASR Models for Handling Multiple Talkers
Chenda Li, Yao Qian, Zhuo Chen 0006, Naoyuki Kanda, Dongmei Wang, Takuya Yoshioka, Yanmin Qian, Michael Zeng 0001 |
INTERSPEECH | 5 |
| 2023 | Speaker Diarization for ASR Output with T-vectors: A Sequence Classification Approach
Midia Yousefi, Naoyuki Kanda, Dongmei Wang, Zhuo Chen 0006, Xiaofei Wang 0009, Takuya Yoshioka |
INTERSPEECH | 3 |
| 2022 | Picknet: Real-Time Channel Selection for Ad Hoc Microphone ArraysabstractThis paper proposes PickNet, a neural network model for real-time channel selection using an ad hoc microphone array. Assuming at most one person to be vocally active at each time point, PickNet identifies the device that is spatially closest to the active person for each time frame by using a short spectral patch of just hundreds of milliseconds. The model is applied to every time frame, and the short time frame signals from the selected microphones are concatenated across the frames to produce an output signal. As the personal devices are usually held close to their owners, the output signal is expected to have higher signal-to-noise and direct-to-reverberation ratios on average than the input signals. Since PickNet utilizes only limited acoustic context at each time frame, the system using the proposed model works in real time and is robust to changes in acoustic conditions. Speech recognition-based evaluation was carried out by using real conversational recordings obtained with various smart-phones. The proposed model yielded significant gains in word error rate with limited computational cost over systems using a block-online beamformer and a single distant microphone. Takuya Yoshioka, Xiaofei Wang 0009, Dongmei Wang |
ICASSP | 3 |
| 2022 | VarArray: Array-Geometry-Agnostic Continuous Speech SeparationabstractContinuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper proposes VarArray, an array-geometry-agnostic speech separation neural network model. The proposed model is applicable to any number of microphones without retraining while leveraging the nonlinear correlation between the input channels. The proposed method adapts different elements that were proposed before separately, including transform-average-concatenate, conformer speech separation, and inter-channel phase differences, and combines them in an efficient and cohesive way. Large-scale evaluation was performed with two real meeting transcription tasks by using a fully developed transcription system requiring no prior knowledge such as reference segmentations, which allowed us to measure the impact that the continuous speech separation system could have in realistic settings. The proposed model outperformed a previous approach to array-geometry-agnostic modeling for all of the geometry configurations considered, achieving asclite-based speaker-agnostic word error rates of 17.5% and 20.4% for the AMI development and evaluation sets, respectively, in the end-to-end setting using no ground-truth segmentations. Takuya Yoshioka, Xiaofei Wang 0009, Dongmei Wang, Zirun Zhu, Zhuo Chen 0006, Naoyuki Kanda |
ICASSP | 3 |
| 2022 | All-Neural Beamformer for Continuous Speech SeparationabstractContinuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common application scenario is transcribing a meeting conversation recorded by a microphone array. Prior studies explored various deep learning models for time-frequency mask estimation, followed by a minimum variance distortionless response (MVDR) filter to improve the automatic speech recognition (ASR) accuracy. The performance of these methods is fundamentally upper-bounded by MVDR’s spatial selectivity. Recently, the all deep learning MVDR (ADL-MVDR) model was proposed for neural beamforming and demonstrated superior performance in a target speech extraction task using pre-segmented input. In this paper, we further adapt ADL-MVDR to the CSS task with several enhancements to enable end-to-end neural beamforming. The proposed system achieves significant word error rate reduction over a baseline spectral masking system on the LibriCSS dataset. Moreover, the proposed neural beamformer is shown to be comparable to a state-of-the-art MVDR-based system in real meeting transcription tasks, including AMI, while showing potentials to further simplify the run-time implementation and reduce the system latency with frame-wise processing. Zhuohuang Zhang, Takuya Yoshioka, Naoyuki Kanda, Zhuo Chen 0006, Xiaofei Wang 0009, Dongmei Wang, Sefik Emre Eskimez |
ICASSP | 6 |
| 2022 | Leveraging Real Conversational Data for Multi-Channel Continuous Speech SeparationabstractExisting multi-channel continuous speech separation (CSS) models are heavily dependent on supervised data - either simulated data which causes data mismatch between the training and real-data testing, or the real transcribed overlapping data, which is difficult to be acquired, hindering further improvements in the conversational/meeting transcription tasks. In this paper, we propose a three-stage training scheme for the CSS model that can leverage both supervised data and extra large-scale unsupervised real-world conversational data. The scheme consists of two conventional training approaches -- pre-training using simulated data and ASR-loss-based training using transcribed data -- and a novel continuous semi-supervised training between the two, in which the CSS model is further trained by using real data based on the teacher-student learning framework. We apply this scheme to an array-geometry-agnostic CSS model, which can use the multi-channel data collected from any microphone array. Large-scale meeting transcription experiments are carried out on both Microsoft internal meeting data and the AMI meeting corpus. The steady improvement by each training stage has been observed, showing the effect of the proposed method that enables leveraging real conversational data for CSS model training. Xiaofei Wang 0009, Dongmei Wang, Naoyuki Kanda, Sefik Emre Eskimez, Takuya Yoshioka |
INTERSPEECH | 2 |
| 2022 | Innate immune memory and its application to artificial immune systems
Dongmei Wang, Hongbin Dong, Chengyu Tan, Zhenhua Xiao, Sai Liu |
J. Supercomput. | 1 |
| 2022 | NKA: a pathogen dose-based natural killer cell algorithm and its application to classification
Dongmei Wang |
J. Supercomput. | 1 |
| 2021 | Simulation research on safety detection of pattern rope jumping motion based on large data backgroundabstractWith the continuous progress of society and the improvement of economic level, more and more people devote their spare time to physical exercise. With the development of the times, figure rope skipping is a new sport which integrates fitness, entertainment, competition and performance. Nowadays, figure rope skipping, as a new fashionable sport, has become a popular fitness sport. In order to give full play to the fitness function of figure rope skipping on the basis of full consideration of safety, a large-data joint prediction model was proposed based on inverse dynamics is proposed to explore the influence of rope skipping on joint position. Introducing the pattern skipping rope into the school physical education curriculum not only makes the students feel the sports side. At the same time, it can also develop the wisdom and potential of students, entertain the body and mind, and cultivate the spirit of unity and cooperation and disappointment. In the future physical education work, we should vigorously develop the national fitness cause, and make the pattern skipping movement become a campus special project to promote its development and promotion on campus. Dongmei Wang |
Connect. Sci. | 1 |
| 2021 | A Safe Zone SMOTE Oversampling Algorithm Used in Earthquake Prediction Based on Extreme Imbalanced Precursor DataabstractEarthquake prediction based on extreme imbalanced precursor data is a challenging task for standard algorithms. Since even if an area is in an earthquake-prone zone, the proportion of days with earthquakes per year is still a minority. The general method is to generate more artificial data for the minority class that is the earthquake occurrence data. But the most popular oversampling methods generate synthetic samples along line segments that join minority class instances, which is not suitable for earthquake precursor data. In this paper, we propose a Safe Zone Synthetic Minority Oversampling Technique (SZ-SMOTE) oversampling method as an enhancement of the SMOTE data generation mechanism. SZ-SMOTE generates synthetic samples with a concentration mechanism in the hyper-sphere area around each selected minority instances. The performance of SZ-SMOTE is compared against no oversampling, SMOTE and its popular modifications adaptive synthetic sampling (ADASYN) and borderline SMOTE (B-SMOTE) on six different classifiers. The experiment results show that the quality of earthquake prediction using SZ-SMOTE as oversampling algorithm significantly outperforms that of using the other oversampling algorithms. Dongmei Wang, Hongbin Dong, Chengyu Tan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2020 | Neural Speech Separation Using Spatially Distributed MicrophonesabstractThis paper proposes a neural network based speech separation method using spatially distributed microphones.Unlike with traditional microphone array settings, neither the number of microphones nor their spatial arrangement is known in advance, which hinders the use of conventional multi-channel speech separation neural networks based on fixed size input.To overcome this, a novel network architecture is proposed that interleaves inter-channel processing layers and temporal processing layers.The inter-channel processing layers apply a selfattention mechanism along the channel dimension to exploit the information obtained with a varying number of microphones.The temporal processing layers are based on a bidirectional long short term memory (BLSTM) model and applied to each channel independently.The proposed network leverages information across time and space by stacking these two kinds of layers alternately.Our network estimates time-frequency (TF) masks for each speaker, which are then used to generate enhanced speech signals either with TF masking or beamforming.Speech recognition experimental results show that the proposed method significantly outperforms baseline multi-channel speech separation systems. Dongmei Wang, Zhuo Chen 0006, Takuya Yoshioka |
INTERSPEECH | 1 |
| 2019 | An Efficient Genetic Algorithm for Active Space Debris Removal PlanningabstractAs the number of debris increases, the space environment becomes crowded and the risk of collision increases. The active debris removal (ADR) technology is an effective measure to suppress the growth of space debris and stabilize the space environment at a safe level. The ADR based a parent spacecraft carrying multi-sub-satellite is one of the most popular methods to de-orbit space debris. The sub-satellites are equipped with different de-orbit devices to capture the debris and move them to lower orbits. One of the key issues have to deal with is that how to plan an efficient strategy for these sub-satellites to visit all the identified debris in one mission at the lowest weighted sum of energy and time. In order to solve the problem, we model the multi-sub-satellite path planning into a multiple traveler salesmen problem (MTSP). Moreover, we propose an improved multi-sub-population parallel genetic algorithm named Route-Break chromosome pairing, it uses an integer coding method and a new crossover-mutation strategy called single chromosome transformation. Each sub-population evolves independently. Finally, we carry out experiments to testify the efficiency of the proposed algorithm. From the experimental results, we can see that the proposed method has a fast convergence speed and a better solution. Dongmei Wang |
CEC | 1 |
| 2017 | Speech Enhancement Based on Harmonic Estimation Combined with MMSE to Improve Speech Intelligibility for Cochlear Implant Recipients
Dongmei Wang, John H. L. Hansen |
INTERSPEECH | 1 |
| 2017 | Robust Harmonic Features for Classification-Based Pitch EstimationabstractPitch estimation in diverse naturalistic audio streams remains a challenge for speech processing and spoken language technology. In this study, we investigate the use of robust harmonic features for classification-based pitch estimation. The proposed pitch estimation algorithm is composed of two stages: pitch candidate generation and target pitch selection. Based on energy intensity and spectral envelope shape, five types of robust harmonic features are proposed to reflect pitch associated harmonic structure. A neural network is adopted for modeling the relationship between input harmonic features and output pitch salience for each specific pitch candidate. In the test stage, each pitch candidate is assessed with an output salience that indicates the potential as a true pitch value, based on its input feature vector processed through the neural network. Finally, according to the temporal continuity of pitch values, pitch contour tracking is performed using a hidden Markov model (HMM), and the Viterbi algorithm is used for HMM decoding. Experimental results show that the proposed algorithm outperforms several state-of-the-art pitch estimation methods in terms of accuracy in both high and low levels of additive noise. Dongmei Wang, Chengzhu Yu, John H. L. Hansen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | F0 estimation for noisy speech by exploring temporal harmonic structures in local time frequency spectrum segmentabstractIn this paper, we propose a noise robust F0 estimation approach by exploring the temporal harmonic structures in local time-frequency (TF) spectrum segment. Since the speech energy is sparsely distributed on the TF plane, the speech harmonic structures occupied in the higher speech energy TF segment are tending to dominate over noise. Thus, we attempt to derive F0 from such high (signal to noise ratio) SNR TF segments rather than full band signal. Our algorithm comprises of two stages: i) F0 candidate estimation for a series of TF segments; ii) F0 tracking based on the acoustic features of each TF segment as well as the F0 temporal continuity constraints. Experimental results show that our approach outperforms the compared methods in terms of F0 estimation accuracy. Dongmei Wang, John H. L. Hansen |
ICASSP | 1 |
| 2014 | Investigation of the relative perceptual importance of temporal envelope and temporal fine structure between tonal and non-tonal languages
Dongmei Wang, James M. Kates, John H. L. Hansen |
INTERSPEECH | 1 |
| 2014 | Noisy speech enhancement based on long term harmonic model to improve speech intelligibility for hearing impaired listenersabstractThis study proposes a speech enhancement algorithm to improve speech intelligibility for hearing impaired listeners in adverse conditions. The proposed algorithm is based on a long term harmonic model, where the harmonics of target speech are more distinguished from noise spectrum interference. Our method consists of two stages: i) Prominent pitch estimation based on long term harmonic feature analysis and neural network classification. ii) Target speech spectrum estimation with pitch information based on long term noise spectrum extraction. The listening experiment with EAS vocoder speech shows that our algorithm is substantially beneficial for cochlear implant recipients to perceive speech in noisy environment in terms of word recognition rate. Dongmei Wang, Philipos C. Loizou, John H. L. Hansen |
INTERSPEECH | 1 |
| 2014 | F0 estimation in noisy speech based on long-term harmonic feature analysis combined with neural network classification
Dongmei Wang, Philipos C. Loizou, John H. L. Hansen |
INTERSPEECH | 1 |
| 2013 | On provisioning diverse circuits in heterogeneous multi-layer optical networks
Dahai Xu, Guangzhi Li, Byrav Ramamurthy, Angela L. Chiu, Dongmei Wang, Robert D. Doverspike |
Comput. Commun. | 5 |
| 2012 | Pitch Estimation Based on Long Frame Harmonic Model and Short Frame Average Correlation Coefficient
Dongmei Wang, Philipos C. Loizou |
INTERSPEECH | 1 |
| 2011 | Cross-layer failure restoration of IP multicast with applications to IPTV
Murat Yuksel, K. K. Ramakrishnan, Robert D. Doverspike, Rakesh K. Sinha, Guangzhi Li, Kostas N. Oikonomou, Dongmei Wang |
Comput. Networks | 7 |
| 2010 | Design of metro Ethernet networksabstractMetropolitan (metro) carriers are deploying next-generation metro Ethernet networks. This deployment will eventually replace traditional private-line services provided by legacy Time Division Multiplexing (TDM) technologies (such as Digital Cross-Connect Systems and SONET rings) with metro Ethernet services. A critical near-term need is the enabling of rich wireless data applications and, thus, begins a massive upgrade of backhaul links from cellular base stations to their Mobile Telephone Switching Offices (MTSOs), where the various voice and data applications will be provided. These base stations are evolving to 3G/4G technology and require metro Ethernet transport and at significantly higher backhaul technology than the prevalent TDM DS1s (1.5 Mb/s). In this paper, we describe a methodology to enable the rapid introduction of metro Ethernet networks, which includes architecture planning, switch location selection, link placement and sizing, as well as customization for cellular network reliability. Our approach uses a combination of many optimization algorithms and has been integrated into a pragmatic tool used by AT&T network planners. Case studies in very large US metro areas have shown that the tool gives cost-ffective solutions consistent with planner expectations and intuition. Dongmei Wang, David F. Lynch, John G. Klincewicz, Guangzhi Li, Robert D. Doverspike, Moshe Segal |
LANMAN | 1 |
| 2009 | Single Channel Music Source Separation based on Harmonic Structure EstimationabstractSingle channel music separation is a useful but difficult problem in audio signal processing field. In this paper a new method is proposed. The method consists of three stages: estimating the harmonic structure of each source in every frame based on iteration of the mixed spectral peaks, clustering the estimated harmonics into the signals they belong to with pitch and formant information, and synthesizing the music source in time domain. Moreover, the method can solve the octave overlapping problem which is a tough one in the single channel source separation area. The experimental results show that our algorithm can separate the mixed signal and obtains a good subjective audio quality. Dongmei Wang, Qinghua Huang |
ISCAS | 1 |
| 2009 | Fast rerouting for IP multicast in managed IPTV networksabstractRecent deployment of IP based multimedia distribution, especially broadcast TV distribution has increased the importance of simple and fast restoration during IP network failures for service providers. In this paper, we propose and evaluate a simple but efficient method for fast rerouting of IP multicast traffic during link failures in managed IPTV networks. More specifically, we devise an algorithm for tuning IP link weights so that the multicast routing path and the unicast routing path between any two routers are failure disjoint, allowing us to use unicast IP encapsulation for undelivered multicast packets during link failures. We demonstrate that, our method can be realized with minor modification to the current multicast routing protocol (PIM-SM). We run our prototype implementation in Emulab which shows our method yields to good performance. Ralf Lübben, Guangzhi Li, Dongmei Wang, Robert D. Doverspike, Xiaoming Fu 0001 |
IWQoS | 3 |
| 2008 | Efficient distributed bandwidth management for MPLS fast reroute
Dongmei Wang, Guangzhi Li |
IEEE/ACM Trans. Netw. | 1 |
| 2007 | IP Backbone Design for Multimedia Distribution: Architecture and PerformanceabstractMultimedia distribution, especially broadcast TV distribution over an IP network requires high bandwidth combined with tight latency and loss constraints, even under failure conditions. Due to the high bandwidth requirements of broadcast TV distribution, use of IP-based multicast to distribute TV content is needed for efficient use of capacity. The protection and restoration mechanisms currently adopted in IP backbones use either IGP re-convergence or some form of fast reroute. The IGP re-convergence mechanism is too slow for real-time multimedia distribution while a drawback of fast reroute is that traffic is re-routed on a link-basis (instead of end-to-end) so there can be traffic overlap during failures. By this we mean traffic passing through the same link along the same direction more than once; this requires more link capacity or it will result in congestion. We propose a routing method that interacts with Fast Reroute and multicast to minimize traffic overlap during failures. We also present an algorithm for link weight setting that avoids traffic overlap due to any single link failure. Performance analysis shows that our methods improve network service availability and significantly reduce the impact of failures. Robert D. Doverspike, Guangzhi Li, K. K. Ramakrishnan, Dongmei Wang |
INFOCOM | 5 |
| 2007 | IGP Weight Setting in Multimedia IP NetworksabstractWith more service providers making considerable investments to roll out multimedia services using IP technology, live TV distribution on IP network is expected to grow impressively over next few years. To make efficient use of IP network infrastructure, service providers use multicast to transmit broadcast video content to the receiving nodes while simultaneously using unicast to transport other services on the same network, such as video on demand (VoD), high-speed Internet (HIS) etc. How to engineer different traffic flows on the same network to avoid traffic congestion becomes a major design issue. Although the service traffic flows are from source to receivers, the IP multicast tree is calculated using IGP shortest path from the receivers to the source (backwards from the flow), while the unicast path is calculated from the source to the receivers (same direction as the flow). To minimize congestion, we take advantage of this property and propose an algorithm to tune IGP link weights (such as OSPF) such that the traffic flows for multicast and unicast don't overlap. This proposal provides a natural way for service providers to use the spare capacity on the opposite direction of multicast traffic flow to roll out additional services in the same IP networks without impacting existing multicast services. Dongmei Wang, Guangzhi Li, Robert D. Doverspike |
INFOCOM | 1 |
| 2006 | Efficient Distributed MPLS P2MP Fast Reroute
Guangzhi Li, Dongmei Wang, Robert D. Doverspike |
INFOCOM | 2 |
| 2005 | A GMPLS based control plane testbed for end-to-end servicesabstractWe describe our implementation choices and experiences gained from developing a GMPLS-based control plane test-bed prototype for end-lo-end services. The prototype includes control plane components of routing and signaling protocols and provides rapid provisioning and restoration of connections across multiple optical network domains. Our experience in implementing the multi-domain signaling showed that it was possible to simply adapt the signaling schemes within the OIF UNI specifications. Our test-bed measurements achieved average 27ms provisioning time and 16.3 ms restoration time for connections crossing three domains. Guangzhi Li, Jennifer Yates, Dongmei Wang, Panagiotis Sebos |
BROADNETS | 3 |
| 2005 | Efficient Distributed Solution for MPLS Fast Reroute
Dongmei Wang, Guangzhi Li |
NETWORKING | 1 |
| 2004 | Detailed Study of IP/ Reconfigurable Optical NetworksabstractIP over reconfigurable optical network architectures have been extensively discussed within the research literature over the past few years. However, although reconfigurable optical networks have been deployed and signaling protocols between IP routers and optical networks have been standardized, large IP backbones are typically deployed using the reconfigurable optical networks. One of the most important criteria in determining whether an IP backbone should be carried over a reconfigurable optical network is economic viability - which necessitates a detailed, accurate economic study of IP backbone over reconfigurable optical network architectures. In this paper, we analyze and explore four IP over optical network architectures for a typical large ISP backbone. In contrast with other published claims, our results suggest that an IP over opaque reconfigurable optical network architecture is not economically attractive with current equipment and IP backbone network design requirements. However, for ISPs also carrying large volumes of transport network private line services, our proposed integrated IP over re-configurable optical network architecture may provide an attractive alternative for providing rapid, cost effective failure recovery. Guangzhi Li, Dongmei Wang, Jennifer Yates, Robert D. Doverspike, Charles R. Kalmanek |
BROADNETS | 2 |
| 2004 | Congestion Control in Resilient Packet RingsabstractCongestion control in ring based packet networks is challenging due to the fact that every node in the network runs both a rate adaptation algorithm, analogous to an endpoint algorithm in other network architectures, and a rate allocation algorithm, analogous to switch-based algorithms in other network architectures. This work describes a congestion control algorithm for IEEE 802.17 resilient packet rings called the enhanced conservative mode algorithm that aims to avoid congestion and achieve a fair rate allocation for fairness eligible traffic in the case of a single bottleneck. We first present analysis to show that existing approaches for RPR congestion control (aggressive and conservative mode) have deficiencies. We present simulation results showing that the proposed enhanced conservative mode congestion control algorithm is a significant improvement. In conjunction with other mechanisms specified in the IEEE 802.17 MAC, the proposed algorithm achieves high utilization on the ring with minimal starvation and oscillations, allows sources to fast start, and provides quality of service for multiple classes of service that require rate, delay and jitter guarantees. Dongmei Wang, K. K. Ramakrishnan, Charles R. Kalmanek, Robert D. Doverspike, Aleksandra Smiljanic |
ICNP | 1 |
| 2004 | Accurate, scalable in-network identification of p2p traffic using application signaturesabstractThe ability to accurately identify the network traffic associated with different P2P applications is important to a broad range of network operations including application-specific traffic engineering, capacity planning, provisioning, service differentiation,etc. However, traditional traffic to higher-level application mapping techniques such as default server TCP or UDP network-port baseddisambiguation is highly inaccurate for some P2P applications.In this paper, we provide an efficient approach for identifying the P2P application traffic through application level signatures. We firstidentify the application level signatures by examining some available documentations, and packet-level traces. We then utilize the identified signatures to develop online filters that can efficiently and accurately track the P2P traffic even on high-speed network links.We examine the performance of our application-level identification approach using five popular P2P protocols. Our measurements show thatour technique achieves less than 5% false positive and false negative ratios in most cases. We also show that our approach only requires the examination of the very first few packets (less than 10packets) to identify a P2P connection, which makes our approach highly scalable. Our technique can significantly improve the P2P traffic volume estimates over what pure network port based approaches provide. For instance, we were able to identify 3 times as much traffic for the popular Kazaa P2P protocol, compared to the traditional port-based approach. Subhabrata Sen, Oliver Spatscheck, Dongmei Wang |
WWW | 3 |
| 2003 | An Efficient Algorithm for OSPF Subnet AggregationabstractMultiple addresses within an OSPF area can be aggregated and advertised together to other areas. This process is known as address aggregation and is used to reduce router computational overheads and memory requirements and to reduce the network bandwidth consumed by OSPF messages. The downside of address aggregation is that it leads to information loss and consequently sub-optimal (non-shortest path) routing of data packets. The resulting difference (path selection error) between the length of the actual forwarding path and the shortest path varies between different sources and destinations. This paper proves that the path selection error from any source to any destination can be bounded using only parameters describing the destination area. Based on this, the paper presents an efficient algorithm that generates the minimum number of aggregates subject to a maximum allowed path selection error. A major operational benefit of our algorithm is that network administrators can select aggregates for an area based solely on the topology of the area without worrying about remaining areas of the OSPF network. The other benefit is that the algorithm enables trade-offs between the number of aggregates and the bound on the path selection error. The paper also evaluates the algorithm's performance on random topologies. Our results show that in some cases, the algorithm is capable of reducing the number of aggregates by as much as 50% with only a relatively small introduction of maximum path selection error. Aman Shaikh, Dongmei Wang, Guangzhi Li, Jennifer Yates, Charles R. Kalmanek |
ICNP | 2 |
| 2003 | Efficient distributed restoration path selection for shared mesh restorationabstractIn MPLS/GMPLS networks, a range of restoration schemes will be required to support different tradeoffs between service interruption time and network resource utilization. In light of these tradeoffs, path-based end-to-end shared mesh restoration provides a very attractive solution. However, efficient use of bandwidth for shared mesh restoration strongly relies on the procedure for selecting restoration paths. We propose an efficient restoration path selection algorithm for restorable connections over shared bandwidth in a fully distributed MPLS/GMPLS architecture. We also describe how to extend MPLS/GMPLS signaling protocols to collect the necessary information efficiently. To evaluate the algorithm's performance, we compare it via simulation with two other well-known algorithms on a typical intercity backbone network. The key figure of merit for restoration bandwidth efficiency is restoration overbuild, i.e., the extra bandwidth required to meet the network restoration objective as a percentage of the bandwidth of the network with no restoration. Our simulation results show that our algorithm uses significantly less restoration overbuild (63%-68%) compared with the other two algorithms (83%-90%). Guangzhi Li, Dongmei Wang, Charles R. Kalmanek, Robert D. Doverspike |
IEEE/ACM Trans. Netw. | 2 |
| 2002 | Efficient Distributed Path Selection for Shared Restoration ConnectionsabstractIn MPLS/GMPLS networks, a range of restoration schemes are required to support different tradeoffs between service interruption time and network resource utilization. In light of these tradeoffs, path-based, end-to-end shared restoration provides a very attractive solution. However, efficient use of capacity for shared restoration strongly relies on the selection procedure of restoration paths. We propose an efficient path-selection algorithm for restoration of connections over shared bandwidth in a fully distributed GMPLS architecture. We also describe how to extend GMPLS signaling protocols to collect the necessary information efficiently. To evaluate the algorithm's performance, we compare it via simulation with two other well-known algorithm on a typical intercity backbone network. The key figure-of-merit for restoration capacity efficiency is restoration overbuild, i.e., the extra capacity required to meet the network restoration objective as a percentage of the capacity of the network with no restoration. Our simulation results show that our algorithm uses significantly less restoration overbuild (63-68%) compared to the other two algorithms (83-90%). Guangzhi Li, Dongmei Wang, Charles R. Kalmanek, Robert D. Doverspike |
INFOCOM | 2 |
| 2001 | A feedback system for graphics video coding and networkingabstractThis paper presents feedback design of a desktop visual communications system. The system consists of video coding using a 3-D graphics model and video transmission over the Internet. To reduce the overall video quality degradation in visual communications caused by coding and networking errors, we jointly researched the compression and transmission of video signals. In building the 3-D graphics model-based coding structure, instead of using feedforward based on pixel intensity, we develop a three-level signal representation and an analysis-by-synthesis feedback framework. In prototyping the video over IP in desktop conferencing, we adapt the video encoders coding rates based on feedback of the network states and receivers. We therefore contribute to the analysis, modeling, and transmission of multimedia signals for desktop visual communications. Dongmei Wang, Russell M. Mersereau |
ICASSP | 1 |
| 2000 | Handling Disaggregate Spatiotemporal Travel Data in GIS
Shih-Lung Shaw, Dongmei Wang |
GeoInformatica | 2 |
| 1996 | Transform predictive coding of wideband speech signalsabstractThis paper presents a novel wideband speech coding algorithm called transform predictive coding (TPC). The main emphasis is on low complexity. TPC uses short-term and long-term prediction to remove the redundancy in speech. The prediction residual is quantized in the frequency domain based on a calculated noise masking threshold. In its simplest form, the TPC coder uses only open-loop quantization and therefore has a low complexity. A 16 kb/s full-duplex, open-loop TPC coder takes only 22% of the CPU load on a 150 MHz SGI Indy workstation and about 34% on a 90 MHz Pentium PC. The speech quality of TPC is almost transparent at 32 kb/s, very good at 24 kb/s, and acceptable at 16 kb/s. In the second half of the paper, we report our recent progress in using closed-loop quantization techniques to improve TPC output speech quality. Juin-Hwey Chen, Dongmei Wang |
ICASSP | 2 |
| 1995 | Codebook adaptation algorithm for a scene adaptive video coderabstractProposes a codebook adaptation algorithm for very low bit rate, real-time video coding. Although adaptive codebook design has been studied in the past, its implementation at very low coding rates suitable for the MPEG4 standard remains significantly challenging. The coder uses a standard motion compensated predictor with DCT quantization. It is unique in that it uses a hybrid scalar/vector quantizer to code predictor residuals. Bits are dynamically allocated to minimize distortion in the current frame, and scalar quantized blocks are used to adapt the VQ codebook. A codebook adaptation algorithm is described which uses an "equidistortion principle" and a competitive learning algorithm to continuously adapt the codewords. This training algorithm results in an increased use of the more efficient vector quantizer and improved video quality. Dongmei Wang, John Hartung |
ICASSP | 1 |