Weibei Dou

dblp:06/6882 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0001-8555-2776ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
abstract
Text-to-audio (TTA) generation with finegrained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works.However, constrained by data scarcity, their generation performance at scale is still compromised.In this study, we recast controllable TTA generation as a multi-task learning problem and introduce a progressive diffusion modeling approach, Con-trolAudio.Our method adeptly fits distributions conditioned on more fine-grained information, including text, timing, and phoneme features, through a step-by-step strategy.First, we propose a data construction method spanning both annotation and simulation, augmenting condition information in the sequence of text, timing, and phoneme.Second, at the model training stage, we pretrain a diffusion transformer (DiT) on large-scale text-audio pairs, achieving scalable TTA generation, and then incrementally integrate the timing and phoneme features with unified semantic representations, expanding controllability.Finally, at the inference stage, we propose progressively guided generation, which sequentially emphasizes more fine-grained information, aligning inherently with the coarse-tofine sampling nature of DiT.Extensive experiments show that ControlAudio achieves stateof-the-art performance in terms of temporal accuracy and speech clarity, significantly outperforming existing methods on both objective and subjective evaluations.Demo samples are available at:
Zehua Chen 0005, Zeqian Ju, Yusheng Dai, Weibei Dou, Jun Zhu 0001
ACL (1)5
2025 BAMN: Brain Asymmetry Analysis Based on Multiplex Networks
abstract
Changes in brain functional asymmetry are important physiological characteristics for evaluating neurorehabilitation. The characteristics of brain networks can be used to assess the brain functional asymmetry. The multiplex networks is defined as a multilayer networks that the interlayer connections are not present, apart from those between replica nodes. How to integrate different single layer networks information to assess the asymmetry for enhancing the assessment accuracy of the neurorehabilitation is a problem that attracts our attention. BAMN, a brain asymmetry analysis method based on multiplex network is presented in this paper. The proposed method extends the attributes of graph theory of single layer to multilayer, calculates their differences between the left and right hemisphere of the brain. It has been validated by using clinical EEG data of after anterior cruciate ligament reconstruction (ACLR) patients and healthy controls, discovering the distinct asymmetry features between these two groups. Some of the signifcant difference features are significantly correlated with the clinical scores and they may be used in the future assessment of the neurorehabilitation.
Xuxin Cai, Lining Zhang, Weibei Dou
BIBM4
2025 AVS3P10 Standard for Real-time Speech Coding
abstract
As the tenth part of the third-generation AVS standard series for real-time speech coding, AVS3P10 is the recent standard completed in the Audio Video Coding Standards Workgroup of China (AVS). Combining the state-of-the-art deep generative networks and signal processing methods, AVS3P10 targets defining new generation neural speech codecs with high quality at low bitrates, enabling excellent experiences even when the bitrate is at 5.9 kbps with excellent error resilience. Moreover, it provides wideband and super wideband coding modes, and it supports the extension of stereo coding. Both subjective listening test and objective measurement prove the merit of AVS3P10. Especially, a lightweight model with only 880k parameters is incorporated to maintain the practicality of AVS3P10 in computational efficiency. Conclusively, AVS3P10 demonstrates the maturity of neural speech coding with broad application perspectives in real-time communication.
Weibei Dou, Gaoxiong Yi, Jingxin Li, Shidong Shang
ICASSP2
2025 FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
abstract
Text-to-audio (T2A) generation has achieved promising results with the recent advances in generative models. However, because of the limited quality and quantity of temporally-aligned audio-text pairs, existing T2A methods struggle to handle the complex text prompts that contain precise timing control, e.g., owl hooted at 2.4s-5.2s. Recent works have explored data augmentation techniques or introduced timing conditions as model inputs to enable timing-conditioned 10-second T2A generation, while their synthesis quality is still limited. In this work, we propose a novel training-free timing-controlled T2A framework, FreeAudio, making the first attempt to enable timing-controlled long-form T2A generation, e.g., owl hooted at 2.4s-5.2s and crickets chirping at 0s-24s. Specifically, we first employ an LLM to plan non-overlapping time windows and recaption each with a refined natural language description, based on the input text and timing prompts. Then we introduce: 1) Decoupling & Aggregating Attention Control for precise timing control; 2) Contextual Latent Composition for local smoothness and Reference Guidance for global consistency. Extensive experiments show that: 1) FreeAudio achieves state-of-the-art timing-conditioned T2A synthesis quality among training-free methods and is comparable to leading training-based methods; 2) FreeAudio demonstrates comparable long-form generation quality with training-based Stable Audio and paves the way for timing-controlled long-form T2A synthesis. Demo samples are available at: https://freeaudio.github.io/FreeAudio/.
Zehua Chen 0005, Zeqian Ju, Weibei Dou, Jun Zhu 0001
ACM Multimedia5
2024 Multi Models Fusion for Predicting Stroke Recovery Based on EEG Features
abstract
Predicting stroke recovery outcomes is crucial yet challenging, and different models yield varied predictions. For addressing the fusion of different models to provide more reliable prediction, a Dempster-Shafer Theory (DST)-based multi-model fusion method for different machine learning models is proposed in this paper to provide a robust prediction for the recovery outcomes. The Shannon entropy is applied as a measure of uncertainty and for performing DST fusion. It is well-suited for multi-model fusion, even though the output distribution of sub-models is unknown. The EEG-based Motor Imagery Brain-Computer Interface (MI-BCI) training is used in the experiments and validations on 13 stroke patients, with the recovery outcomes prediction accuracy of 92.3% and uncertainty of 0.24 of the multi-model fusion, demonstrating higher accuracy and lower uncertainty compared to any single prediction model.
Mingchen Cheng, Xueyi Ni, Zinan Yuan, Zexuan Hao, Zhimin Shao, Weibei Dou
BIBM8
2023 Enhancing Artifact Removal From Scalp EEG Using State-Wise Deep Convolutional Network
abstract
Pre-processing is a fundamental step for any tasks reliant on scalp EEG data. The presence of various artifacts in acquired EEG data, which mask the expected features of brain activity, underscores the pivotal role of artifact removal in pre-processing. Recently, deep learning methods have demonstrated superior efficacy in artifact removal compared to conventional methods such as regression, classical filtering, and signal decomposition. While EEG signals, characterized as time series, differ from extensively explored modalities like photos or videos in multimodal machine learning, a suitable network architecture is necessary for the utilization of deep learning methods in EEG artifact removal tasks. Therefore, in this study, we propose a neural network architecture utilizing self-learned state distinction criteria for time series segmentation and tested its artifact removal performance on semi-simulated and real EEG datasets. Notably, our proposed network architecture outperforms the state-of-the-art network in artifact removal performance. Furthermore, the proposed model demonstrates acceptable real-time processing capabilities, thus highlighting its potential applications in real clinical and research settings as an online pre-processing step.
Xuepeng Huang, Zexuan Hao, Weibei Dou
BIBM4
2022 Automatic Removal of Scalp EEG Artifacts Using an Interpretable Hybrid Deep Learning Method
abstract
The application of scalp electroencephalogram (EEG) in brain-computer interface (BCI) is inevitably hindered by the contamination of various physiological artifacts. Artifact elimination plays an important role in EEG signal analysis due to its weak power and low signal-to-noise ratio. Recent studies using deep learning (DL) methods have reported good performance for offline artifact elimination. However, existing methods still neglect two key defects in clinical applications: (1) offline processing may not fit inreal-time BCI application scenarios; and (2) direct and arbitrary usage of conventional neural network structures lacks physiological interpretability. To address these problems, we propose a hybrid method for EEG artifact elimination, combining a canonical correlation analysis (CCA) module–a blind source separation method with a multi-scaled convolutional neural network (CNN) module. Specifically, the CCA module transfers the recorded EEG signals simultaneously from the channel space to the source space and the CNN module detects and eliminates artifacts with explainable processing. Extensive experiments on semi-simulation, resting-state, and motor imagery EEG datasets were conducted and the results showed competitive performance and strong adaptability of the proposed method.
Chenhao Bao, Zexuan Hao, Weibei Dou
BIBM3
2022 Clique Network-Based Statistics for Detecting Altered Topological Structures in the Brain Network
abstract
Networks can model structural and functional connection of human brain. Clique as a fully connected structure reveals valuable interaction pattern of network. Some statistical methods, such as network-based statistics (NBS), are frequently utilized to detect discrepancy connections and connected components of networks. The existing NBS family methods focus on the identification of different expressed connected component and fail to extract special topological structure. In this paper, a clique network-based statistics (clique-NBS) method is proposed as an extension of NBS to detect discrepant cliques of networks, and its effect has been validated by both simulated and clinical data. The clinical data is the resting state BOLD-fMRI of 29 multiple system atrophy (MSA) patients and 27 healthy controls for detecting altered functional modules after MSA. The proposed methods successfully identified 28 altered functional modules containing cerebellar parts. By comparing with edge-wise FDR correction method, the proposed method shows an increased performance on true positive rate. As a result, the proposed clique-NBS extends the existing NBS family, and it can be used to find and investigate the discrepant subnetworks with clique topological structure in human brain.
Yunxiang Ge, Weibei Dou
BIBM3
2022 MES-P: An Emotional Tonal Speech Dataset in Mandarin with Distal and Proximal Labels
abstract
Emotion shapes all aspects of our interpersonal and intellectual experiences. Its automatic analysis has therefore many applications. In this paper, we propose an emotional tonal speech dataset, Mandarin Chinese Emotional Speech Dataset-Portrayed (MES-P), with both distal and proximal labels. In contrast with state of the art datasets which only focused on perceived emotions, MES-P includes not only perceived emotions (proximal labels) but also intended emotions (distal labels), to make it possible to study human emotional intelligence, i.e., emotion expression/understanding ability, and emotional misunderstandings in real life. Furthermore, MES-P also captures a main feature of tonal languages, and provides emotional speech samples matching the tonal distribution in real life Mandarin. MES-P dataset also features emotion intensity variations, by introducing both moderate and intense versions for joy, anger, and sadness, in addition to neutral. Ratings of the collected speech samples are made in valence-arousal space through continuous coordinate locations, resulting in an emotional distribution pattern in 2D VA space. High consistency between the speakers emotional intentions and the listeners perceptions is also proved by Cohens Kappa coefficients. Finally, extensive experiments are carried out as a baseline on MES-P for automatic emotion recognition and with comparison to human emotion intelligence.
Zhongzhe Xiao, Weibei Dou, Liming Chen 0002
IEEE Trans. Affect. Comput.3
2021 Extended Network-Based Statistics for Measuring Altered Directed Connectivity Components in the Human Brain
abstract
The human brain is thought to be an interconnected network. In network analysis of the human brain, one of the most concerned issue is how to measure the altered directed connectivity component. In this paper, extended network-based statistics (e-NBS) is proposed to measure the directed connectivity component alteration. The proposed method can identify weakly and strongly connected components of directed networks using t-test and non-parametric test under a range of threshold values. A directed brain network construction method using convergent cross-mapping (CCM) is also proposed, which estimates causal relationships from the nonlinear state space reconstruction of time series of neural signal. We validated the proposed methods using resting-state BOLD fMRI data of 23 SCI patients and 22 healthy controls. The proposed methods showed significant different connected components under several e-NBS parameter settings. A common stable connected component pattern was also identified. The proposed method provides a new way to investigate altered directed connectivity components in the human brain.
Yunxiang Ge, Yutong Feng, Weibei Dou
BIBM5
2021 Modified Linear Fascicle Evaluation (mLiFE) for Improving the Fiber Tractography of Stroke Patients using Diffusion MRI
abstract
Tractography based on DTI (Diffusion Tensor Imaging) data can noninvasively probe the white matter structure, thus helps the assessment of central nerve injury such as stroke. The Linear Fascicle Evaluation (LiFE) can compare tractography algorithms by evaluating the tractography results. However, considering the pathological changes of stroke patients, the result in lesion area is doubtful. A modified LiFE (mLiFE) algorithm is proposed in this paper to improve the evaluation accuracy of the tractography results in lesion areas. mLiFE assumes that the diffusion of voxels without fiber crossing is isotropic and computes the main diffusion eigenvalue for every voxel, instead of choosing a fixed value as in LiFE. Additionally, by applying mLiFE in the tractography method which uses the Sparse fascicle models (SFM), the modified SFM (mSFM) is also proposed in this paper. Experiments on the DTI data of stroke patients who received neural stem cell transplantation prove that the proposed mLiFE can evaluate tractography results in the lesion area. In addition, mSFM can reflect neural fiber recovery following surgery efficiently. The proposed mLiFE enables accurate evaluation and fiber estimation for stroke patients and could provide insights of neural fiber recovery for clinicians
Yunxiang Ge, Weibei Dou, Guangzhu Zhang 0002
BIBM3
2016 A biologically plausible spiking model for interaural level difference processing auditory pathway in human brain
abstract
Human brain provides an accurate and energy efficient function for localizing sound sources. In order to build a localizing system with similar advantages, a biologically plausible spiking model for interaural level difference (ILD) processing auditory pathway from cochlea to the inferior colliclus in human brain is proposed in this paper. The biological plausibility is based on the facts that all nodes in the model are implemented as the Integrate-and-Fire neurons, it does not utilize complex mathematical functions, and the architecture bears a high resemblance to the physiological structure. A one-shot learning algorithm characterized as simple and fast is designed to construct the synapse weights between the lateral superior olive and the inferior colliclus in order to render the proposed model to produce expected responses for different azimuthal angles of sounds. Some simulation results are presented to demonstrate the abilities of the proposed model to identify the angle of the input stereo sound and to distinguish multiple sounds coming from different angles. It is also shown that the proposed model behaves in the way that is in high accordance with former physiological and psychoacoustic conclusions.
Weibei Dou
IJCNN2
2014 Facial expression recognition and generation using sparse autoencoder
abstract
Facial expression recognition has important practical applications. In this paper, we propose a method based on the combination of optical flow and a deep neural network—stacked sparse autoencoder (SAE). This method classifies facial expressions into six categories (i.e. happiness, sadness, anger, fear, disgust and surprise). In order to extract the representation of facial expressions, we choose the optical flow method because it could analyze video image sequences effectively and reduce the influence of personal appearance difference on facial expression recognition. Then, we train the stacked SAE with the optical flow field as the input to extract high-level features. To achieve classification, we apply a softmax classifier on the top layer of the stacked SAE. This method is applied to the Extended Cohn-Kanade Dataset (CK+). The expression classification result shows that the SAE performances the classification effectively and successfully. Further experiments (transformation and purification) are carried out to illustrate the application of the feature extraction and input reconstruction ability of SAE.
Xueshi Hou, Guangda Su, Weibei Dou
SMARTCOMP6
2013 MDCT Sinusoidal Analysis for Audio Signals Analysis and Processing
abstract
The Modified Discrete Cosine Transform (MDCT) is widely used in audio signals compression, but mostly limited to representing audio signals. This is because the MDCT is a real transform: Phase information is missing and spectral power varies frame to frame even for pure sine waves. We have a key observation concerning the structure of the MDCT spectrum of a sine wave: Across frames, the complete spectrum changes substantially, but if separated into even and odd subspectra, neither changes except scaling. Inspired by this observation, we find that the MDCT spectrum of a sine wave can be represented as an envelope factor times a phase-modulation factor. The first one is shift-invariant and depends only on the sine wave's amplitude and frequency, thus stays constant over frames. The second one has the form of sinθ for all odd bins and cosθ for all even bins, leading to subspectra's constant shapes. But this θ depends on the start point of a transform frame, therefore, changes at each new frame, and then changes the whole spectrum. We apply this formulation of the MDCT spectral structure to frequency estimation in the MDCT domain, both for pure sine waves and sine waves with noises. Compared to existing methods, ours are more accurate and more general (not limited to the sine window). We also apply the spectral structure to stereo coding. A pure tone or tone-dominant stereo signal may have very different left and right MDCT spectra, but their subspectra have similar shapes. One ratio for even bins and one ratio for odd bins will be enough to reconstruct the right from the left, saving half bitrate. This scheme is simple and at the same time more efficient than the traditional Intensity Stereo (IS).
Weibei Dou, Huazhong Yang
IEEE Trans. Speech Audio Process.2
2011 DFT spectrum estimation from critically sampled lapped transforms
Weibei Dou, Huazhong Yang
Signal Process.2
2010 MDCT spectrum separation: Catching the fine spectral structures for stereo coding
abstract
The spectrum of a sinusoid using the Modified Discrete Cosine Transform (MDCT), when separated into an even subspectrum and an odd subspectrum by bin parity, gives rise to a distinctive property-subspectral shapes are independent of the sinusoid phase, which contributes only to scaling. Based on this finding, we propose an Even-Odd (EO) scheme for stereo coding: partitioning the even and odd subspectra separately into subbands to capture the fine spectral structures of sinusoidal and rich tone signals. The scheme reduces the coding noises by 0-20 dB for music signals. When integrated into a MDCT domain KLT-based stereo coder, the scheme boosts subjective listening test (MUSHRA) scores. This coder, called KLT-EO, competes the Parametric Stereo (PS) in quality by a slightly higher bitrate but without the algorithmic delay of 20 ms resulted from the stereo processing.
Weibei Dou, Ping Chi, Huazhong Yang
ICASSP2
2010 Maximal Coherence Rotation for stereo coding
abstract
This paper presents a linear operation called Maximal Coherence Rotation (MCR) on paired vectors for stereo audio coding. Intrigued by the idea that stronger coherence between paired channels will lead to higher stereo coding efficiency, we develop MCR to maximize the coherence restricted by being invertible and energy-conserving. It results in equal energy, minimized difference, and always non-negative coherence for the pair of channels processed. In binaural hearing, this can be viewed as turning a physical sound source at any azimuth to a virtual one on the median plane. A prototype MCR stereo coder shows significantly higher quality for some test sequences than that of AMR-WB+, one of the best low bitrate stereo coders in the public domain. And as a preprocessing tool for MPEG-4 Parametric Stereo (PS), MCR avoids out-of-phase inter-channel cancellation during 2-to-1 channel downmixing without any additional bandwidth requirement, thanks to the maximized coherence.
Weibei Dou, Huazhong Yang
ICME2
2010 Multi-stage classification of emotional speech motivated by a dimensional emotion model
Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002
Multim. Tools Appl.3
2009 Image Categorization Using ESFS: A New Embedded Feature Selection Method Based on SFS
Huanzhang Fu, Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002
ACIVS4
2007 Fuzzy kappa for the agreement measure of fuzzy classifications
Weibei Dou, Su Ruan, Yanping Chen 0003, Daniel Bloyet, Jean-Marc Constans
Neurocomputing1
2007 A framework of fuzzy information fusion for the segmentation of brain tumor tissues on MR images
Weibei Dou, Su Ruan, Yanping Chen 0003, Daniel Bloyet, Jean-Marc Constans
Image Vis. Comput.1
2006 Combining Short and Long Term Audio Features for TV Sports Highlight Detection
Weibei Dou, Liming Chen 0002
ECIR2
2005 Features extraction and selection for emotional speech classification
abstract
The classification of emotional speech is a topic in speech recognition with more and more interest, and it has giant prospect in applications in a wide variety of fields. It is an important preparation for automatic classification and recognition of emotions to select a proper feature set as a description to the emotional speech, and to find a proper definition to the emotions in speech. The speech samples used in this paper come from Berlin database which contains 7 kinds of emotions, with 207 speech samples of male voice and 287 speech samples of female voice. A feature set of 50 potentially features is extracted and analyzed, and the best features are selected. A definition of emotions as 3-states emotions is also proposed in this paper.
Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002
AVSS3
2004 Possibilistic-clustering-based MR brain image segmentation with accurate initialization
abstract
Magnetic resonance image analysis by computer is useful to aid diagnosis of malady. We present in this paper a automatic segmentation method for principal brain tissues. It is based on the possibilistic clustering approach, which is an improved fuzzy c-means clustering method. In order to improve the efficiency of clustering process, the initial value problem is discussed and solved by combining with a histogram analysis method. Our method can automatically determine number of classes to cluster and the initial values for each class. It has been tested on a set of forty MR brain images with or without the presence of tumor. The experimental results showed that it is simple, rapid and robust to segment the principal brain tissues.
Qingmin Liao, Yingying Deng, Weibei Dou, Su Ruan, Daniel Bloyet
VCIP3
2001 New window-switching criterion of audio compression
abstract
The ERIC (energy ramp in critical-band) criterion for window switching in perceptual audio coding is proposed. This new criterion can effectively distinguish real burst signals from stationary signals with high perceptual entropy. Therefore, comparing with the perceptual entropy criterion used in the MPEG standard, the possibility of misjudgment of the window types is reduced. It results in some improvements on both the audio quality and the coding efficiency.
Zhaorong Hou, Weibei Dou, Zaiwang Dong
MMSP2