VLDB 2026 Research / reviewers in the wild / expert
Heming Wang
dblp:146/1559
· DBLP profile ↗
27ranked-venue papers
13as first author
23since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Talker-independent multi-pitch tracking in noisy and reverberant scenarios
Heming Wang |
Comput. Speech Lang. | 2 |
| 2026 | Adaptive disentangled learning recommendation via similarity popularity
Jianmei Ye, Heming Wang, Jiangzhou Deng, Yong Wang 0009, Zeshui Xu, Kobiljon Kh. Khushvakhtzoda |
Appl. Intell. | 2 |
| 2026 | A speech prediction model based on codec modeling and transformer decodingabstractSpeech prediction is essential for tasks like packet loss concealment and algorithmic delay compensation. This paper proposes a novel prediction algorithm that leverages a speech codec and transformer decoder to autoregressively predict missing frames. Unlike text-guided methods requiring auxiliary information, the proposed approach operates solely on speech for prediction. A comparative study is conducted to evaluate and compare the proposed and existing speech prediction methods on packet loss concealment (PLC) and frame-wise speech prediction tasks. Comprehensive experiments demonstrate that the proposed model achieves superior prediction results, which are substantially better than other state-of-the-art baselines, including on a recent PLC challenge. We also systematically examine factors influencing prediction performance, including context window lengths, prediction lengths, and training and inference strategies. • We propose a codec-based speech prediction approach that effectively leverages acoustic tokens and embeddings extracted from speech codecs. • We systematically evaluate the speech prediction performance of the proposed approach, and demonstrate that it outperforms recent baselines. • We investigate the factors that impact speech prediction performance and examine different training and inference strategies. Heming Wang |
Comput. Speech Lang. | 1 |
| 2026 | Joint robust transmit waveform and receive beamforming design for MIMO dual-function radar-communication systems
Xuchen Liu 0002, Yongjun Liu 0002, Guisheng Liao, Heming Wang, Jiaguo Lu |
Signal Process. | 5 |
| 2026 | Robust waveform design for distributed MIMO dual-function radar-communication systems
Yongjun Liu 0002, Guisheng Liao, Xuchen Liu 0002, Xiaoyang Dong 0005, Heming Wang |
Signal Process. | 6 |
| 2025 | LLMs Empowered MIMO Antenna Decoupling for Wireless Networks: From Mathematical Tuning to Physically-Aware DesignabstractWireless networks increasingly rely on multiple-input-multiple-output (MIMO) antennas to achieve high capacity but often suffer from poor mutual coupling, especially in compact form. Traditional decoupling design approaches struggle to generalize across diverse deployment scenarios. Furthermore, MIMO decoupling workflows often depend on mathematical or manual tuning, which makes them inefficient and inflexible in facing complex electromagnetic (EM) interactions in compact arrays. Although large language models (LLMs) have shown strong generalization and reasoning capabilities across modalities such as images, text, and code, their application to antenna design remains underexplored due to challenges like embedding physical priors, interpreting field data, and integrating with simulation-driven workflows. In this paper, to address the above issues, we propose a physically-aware LLMs empowered scheme to achieve MIMO antenna decoupling. Our method decomposes the decoupling task into three coordinated modules—Analysis, Design, and Optimization, each empowered by multiple specialized LLMs agents. Through joint guidance of geometric constraints with EM field and surface current data given by EM simulation software, our system coordinates multiple multimodal agents to iteratively design and optimize the MIMO antenna configuration. The simulation results indicate that our approach successfully and efficiently improves isolation performance, underscoring the potential of collaborative intelligence in wireless hardware optimization. Heming Wang, Guozhi Hao |
GLOBECOM | 1 |
| 2025 | Combined generative and predictive modeling for speech super-resolution
Heming Wang, Eric W. Healy, DeLiang Wang |
Comput. Speech Lang. | 1 |
| 2025 | A systematic study of DNN based speech enhancement in reverberant and reverberant-noisy environments
Heming Wang, Ashutosh Pandey 0004, DeLiang Wang |
Comput. Speech Lang. | 1 |
| 2024 | SPATIALCODEC: Neural Spatial Speech CodingabstractIn this work, we address the challenge of encoding speech captured by a microphone array using deep learning techniques with the aim of preserving and accurately reconstructing crucial spatial cues embedded in multi-channel recordings. We propose a neural spatial audio coding framework that achieves a high compression ratio, leveraging single-channel neural sub-band codec and SpatialCodec. Our approach encompasses two phases: (i) a neural sub-band codec is designed to encode the reference channel with low bit rates, and (ii), a SpatialCodec captures relative spatial information for accurate multi-channel reconstruction at the decoder end. In addition, we also propose novel evaluation metrics to assess the spatial cue preservation: (i) spatial similarity, which calculates cosine similarity on a spatially intuitive beamspace, and (ii), beamformed audio quality. Our system shows superior spatial performance compared with high bitrate baselines and black-box neural architecture. Demos are available at https://xzwy.github.io/SpatialCodecDemo. Codes and models are available at https://github.com/XZWY/SpatialCodec. Zhongweiyang Xu, Yong Xu 0004, Vinay Kothapally, Heming Wang, Muqiao Yang, Dong Yu 0001 |
ICASSP | 4 |
| 2024 | uSee: Unified Speech Enhancement And Editing with Conditional Diffusion ModelsabstractSpeech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional diffusion models to handle various tasks at the same time in a generative manner. Specifically, by providing multiple types of conditions including self-supervised learning embeddings and proper text prompts to the score-based diffusion model, we can enable controllable generation of the unified speech enhancement and editing model to perform corresponding actions on the source speech. Our experiments show that our proposed uSee model can achieve superior performance in both speech denoising and dereverberation compared to other related generative speech enhancement models, and can perform speech editing given desired environmental sound text description, signal-to-noise ratios (SNR), and room impulse responses (RIR). Demos of the generated speech are available at https://muqiaoy.github.io/usee. Muqiao Yang, Yong Xu 0004, Zhongweiyang Xu, Heming Wang, Bhiksha Raj, Dong Yu 0001 |
ICASSP | 5 |
| 2024 | DDTSE: Discriminative Diffusion Model for Target Speech ExtractionabstractDiffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction under multispeaker noisy conditions remains relatively unexplored. Moreover, the superior quality of diffusion methods typically comes at the cost of slower inference speed. In this paper, we introduce the Discriminative Diffusion model for Target Speech Extraction (DDTSE). We apply the same forward process as diffusion models and utilize the reconstruction loss similar to discriminative methods. Furthermore, we devise a two-stage training strategy to emulate the inference process during model training. DDTSE not only works as a standalone system, but also can further improve the performance of discriminative models without additional retraining. Experimental results demonstrate that DDTSE not only achieves higher perceptual quality but also accelerates the inference process by 3 times compared to the conventional diffusion model. Leying Zhang, Yao Qian, Linfeng Yu, Heming Wang, Hemin Yang, Shujie Liu 0001, Yanmin Qian |
SLT | 4 |
| 2023 | DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation TasksabstractSelf-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem, in this paper, we propose data2vec-SG (Speech Generation), which is a teacher-student learning framework that addresses speech generation tasks. Our data2vec-SG introduces a reconstruction module into data2vec [1] and enforces the representations to contain not only the semantic information but also the acoustic knowledge to generate clean speech waveforms. Experimental results demonstrate that the proposed framework boosts the performance of various speech generation tasks including speech enhancement, speech separation, and packet loss concealment. Meanwhile, the learned representation is also capable of helping other downstream tasks, which is demonstrated by the good performance in the speech recognition task in both clean and noisy conditions. Heming Wang, Yao Qian, Hemin Yang, Naoyuki Kanda, Takuya Yoshioka, Xiaofei Wang 0009, Shujie Liu 0001, Zhuo Chen 0006, DeLiang Wang, Michael Zeng 0001 |
ICASSP | 1 |
| 2023 | Cross-Domain Diffusion Based Speech Enhancement for Very Noisy SpeechabstractDeep learning based speech enhancement has achieved remarkable success, but challenges remain in low signal-to-noise ratio (SNR) nonstationary noise scenarios. In this study, we propose to incorporate diffusion-based learning into an enhancement model and improve robustness in extremely noisy conditions. Specifically, a frequency-domain diffusion-based generative module is employed, and it accepts the enhanced signal obtained from a time-domain supervised enhancement module as an auxiliary input to learn to recover clean speech spectrograms. Experimental results on the TIMIT dataset demonstrate the advantage of this approach and show better enhancement performance over other strong baselines in both -5 and -10 dB SNR noisy conditions. Heming Wang, DeLiang Wang |
ICASSP | 1 |
| 2023 | $F0$ Estimation and Voicing Detection With Cascade Architecture in Noisy SpeechabstractAs a fundamental problem in speech processing, pitch tracking has been studied for decades. While strong performance has been achieved on clean speech, pitch tracking in noisy speech is still challenging. Severe non-stationary noises not only corrupt the harmonic structure in voiced intervals but also make it difficult to determine the existence of voiced speech. Given the importance of voicing detection for pitch tracking, this study proposes a neural cascade architecture that jointly performs pitch estimation and voicing detection. The cascade architecture optimizes a speech enhancement module and a pitch tracking module, and is trained in a speaker-independent and noise-independent way. It is observed that incorporating the enhancement module improves both pitch estimation and voicing detection accuracy, especially in low signal-to-noise ratio (SNR) conditions. In addition, compared with frameworks that combine corresponding single-task models, the proposed multi-task framework achieves better performance and is more efficient. Experimental results show that the proposed method is robust to different noise conditions and substantially outperforms other competitive pitch tracking methods. Yixuan Zhang 0005, Heming Wang, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Wav2vec-Switch: Contrastive Learning from Original-Noisy Speech Pairs for Robust Speech RecognitionabstractThe goal of self-supervised learning (SSL) for automatic speech recognition (ASR) is to learn good speech representations from a large amount of unlabeled speech for the downstream ASR task. However, most SSL frameworks do not consider noise robustness which is crucial for real-world applications. In this paper we propose wav2vec-Switch, a method to encode noise robustness into contextualized representations of speech via contrastive learning. Specifically, we feed original-noisy speech pairs simultaneously into the wav2vec 2.0 network. In addition to the existing contrastive learning task, we switch the quantized representations of the original and noisy speech as additional prediction targets of each other. By doing this, it enforces the network to have consistent predictions for the original and noisy speech, thus allows to learn contextualized representation with noise robustness. Our experiments on synthe-sized and real noisy data show the effectiveness of our method: it achieves 2.9–4.9% relative word error rate (WER) reduction on the synthesized noisy LibriSpeech data without deterioration on the original data, and 5.7% on CHiME-4 real 1-channel noisy data compared to a data augmentation baseline even with a strong language model for decoding. Our results on CHiME-4 can match or even surpass those with well-designed speech enhancement components. Jinyu Li 0001, Heming Wang, Yao Qian, Chengyi Wang 0002, Yu Wu 0012 |
ICASSP | 3 |
| 2022 | Improving Noise Robustness of Contrastive Speech Representation Learning with Speech ReconstructionabstractNoise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this work, instead of suppressing background noise with a conventional cascaded pipeline, we employ a noise-robust representation learned by a refined self-supervised framework for noisy speech recognition. We propose to combine a reconstruction module with contrastive learning and perform multi-task continual pre-training on noisy data. The reconstruction module is used for auxiliary learning to improve the noise robustness of the learned representation and thus is not required during inference. Experiments demonstrate the effectiveness of our proposed method. Our model substantially reduces the word error rate (WER) for the synthesized noisy LibriSpeech test sets, and yields around 4.1/7.5% WER reduction on noisy clean/other test sets compared to data augmentation. For the real-world noisy speech from the CHiME-4 challenge (1-channel track), we have obtained the state of the art ASR performance without any denoising front-end. Moreover, we achieve comparable performance to the best supervised approach reported with only 16% of labeled data. Heming Wang, Yao Qian, Xiaofei Wang 0009, Chengyi Wang 0002, Shujie Liu 0001, Takuya Yoshioka, Jinyu Li 0001, DeLiang Wang |
ICASSP | 1 |
| 2022 | Cross-Domain Speech Enhancement with a Neural Cascade ArchitectureabstractThis paper proposes a novel cascade architecture to address the monaural speech enhancement problem. We leverage three different domains of speech representation, namely spectral magnitude, waveform, and complex spectrogram, to progressively suppress the background noise within noisy speech. Our proposed neural cascade architecture consists of three modules, and each operates on the original noisy input and the output of the previous module in a distinct speech representation. During training, the network simultaneously optimizes all modules with a triple-domain loss. Experiments on the WSJ0 SI-84 corpus demonstrate that our proposed approach achieves superior enhancement results, and substantially outperforms previous baselines in terms of both speech quality and intelligibility. Heming Wang, DeLiang Wang |
ICASSP | 1 |
| 2022 | Attention-Based Fusion for Bone-Conducted and Air-Conducted Speech Enhancement in the Complex DomainabstractBone-conduction (BC) microphones capture speech signals by converting the vibrations of the human skull into electrical signals. BC sensors are insensitive to acoustic noise, but limited in bandwidth. On the other hand, conventional or air-conduction (AC) microphones are capable of capturing full-band speech, but are susceptible to background noise. We propose to combine the strengths of AC and BC microphones by employing a convolutional recurrent network that performs complex spectral mapping. To better utilize signals from both kinds of microphone, we employ attention-based fusion with early-fusion and late-fusion strategies. Experiments demonstrate the superiority of the proposed method over other recent speech enhancement methods combining BC and AC signals. In addition, our enhancement performance is significantly better than conventional speech enhancement counterparts, especially in low signal-to-noise ratio scenarios. Heming Wang, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 1 |
| 2022 | Densely-connected Convolutional Recurrent Network for Fundamental Frequency Estimation in Noisy Speechabstractin turn. Experimental results show that the cascade model brings further improvements to the DC-CRN model, especially in low signal-to-noise ratio (SNR) conditions. Yixuan Zhang 0005, Heming Wang, DeLiang Wang |
INTERSPEECH | 2 |
| 2022 | Large-Scale Worst-Case Topology OptimizationabstractAbstract We propose a novel topology optimization method to efficiently minimize the maximum compliance for a high‐resolution model bearing uncertain external loads. Central to this approach is a modified power method that can quickly compute the maximum eigenvalue to evaluate the worst‐case compliance, enabling our method to be suitable for large‐scale topology optimization. After obtaining the worst‐case compliance, we use the adjoint variable method to perform the sensitivity analysis for updating the density variables. By iteratively computing the worst‐case compliance, performing the sensitivity analysis, and updating the density variables, our algorithm achieves the optimized models with high efficiency. The capability and feasibility of our approach are demonstrated over various large‐scale models. Typically, for a model of size 512×170×170 and 69934 loading nodes, our method took about 50 minutes on a desktop computer with an NVIDIA GTX 1080Ti graphics card with 11 GB memory. Di Zhang 0013, Xiaoya Zhai, Xiao-Ming Fu 0001, Heming Wang, Ligang Liu 0001 |
Comput. Graph. Forum | 4 |
| 2022 | Neural Cascade Architecture With Triple-Domain Loss for Speech EnhancementabstractThis paper proposes a neural cascade architecture to address the monaural speech enhancement problem. The cascade architecture is composed of three modules which optimize in turn enhanced speech with respect to the magnitude spectrogram, the time-domain signal and the complex spectrogram. Each module takes as input the noisy speech and the output obtained from the previous module, and generates a prediction of the respective target. Our model is trained in an end-to-end manner, using a triple-domain loss function that accounts for three domains of signal representation. Experimental results on the WSJ0 SI-84 corpus show that the proposed model outperforms other strong speech enhancement baselines in terms of objective speech quality and intelligibility. Heming Wang, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Fusing Bone-Conduction and Air-Conduction Sensors for Complex-Domain Speech EnhancementabstractSpeech enhancement aims to improve the listening quality and intelligibility of noisy speech in adverse environments. It proves to be challenging to perform speech enhancement in very low signal-to-noise ratio (SNR) conditions. Conventional speech enhancement utilizes air-conduction (AC) microphones, which are sensitive to background noise but capable of capturing full-band signals. On the other hand, bone-conduction (BC) sensors are unaffected by acoustic noise, but recorded speech has limited bandwidth. This study proposes an attention-based fusion method to combine the strengths of AC and BC signals and perform complex spectral mapping for speech enhancement. Experiments on the EMSB dataset demonstrate that the proposed approach effectively leverages the advantages of AC and BC sensors, and outperforms a recent time-domain baseline in all conditions. We also show that the sensor fusion method is superior to single-sensor counterparts, especially in low SNR conditions. As the amount of BC data is very limited, we additionally propose a semi-supervised technique to utilize both parallelly and unparallely recorded AC and BC speech signals. With additional AC speech from the AISHELL-1 dataset, we achieve similar performance to supervised learning with only 50% parallel data. Heming Wang, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Towards Robust Speech Super-ResolutionabstractSpeech super-resolution (SR) aims to increase the sampling rate of a given speech signal by generating high-frequency components. This paper proposes a convolutional neural network (CNN) based SR model that takes advantage of information from both time and frequency domains. Specifically, the proposed CNN is a time-domain model that takes the raw waveform of low-resolution speech as the input, and outputs an estimate of the corresponding high-resolution waveform. During the training stage, we employ a cross-domain loss to optimize the network. We compare our model with several deep neural network (DNN) based SR models, and experiments show that our model outperforms existing models. Furthermore, the robustness of DNN-based models is investigated, in particular regarding microphone channels and downsampling schemes, which have a major impact on the performance of DNN-based SR models. By training with proper datasets and preprocessing, we improve the generalization capability for untrained microphone channels and unknown downsampling schemes. Heming Wang, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Time-Frequency Loss for CNN Based Speech Super-ResolutionabstractSpeech super-resolution (SR), also called speech bandwidth extension (BWE), aims to increase the sampling rate of a given lower resolution speech signal. Recent years have witnessed the successful application of deep neural networks in time or frequency domains, and deep learning has improved the performance considerably compared with conventional approaches. This paper proposes an autoencoder based fully convolutional neural network (CNN) that merges the information from both time and frequency domains. At the training time, we optimize the CNN using a new time-frequency loss (T-F loss), which combines a time domain loss and a frequency domain loss. The experimental results show that our model trained with the T-F loss achieves significantly better results than other state-of-the-art models, and yields balanced performance in terms of time and frequency metrics. Heming Wang, DeLiang Wang |
ICASSP | 1 |
| 2019 | Large-scale prediction of ADAR-mediated effective human A-to-I RNA editingabstractAdenosine-to-inosine (A-to-I) editing by adenosine deaminase acting on the RNA (ADAR) proteins is one of the most frequent modifications during post- and co-transcription. To facilitate the assignment of biological functions to specific editing sites, we designed an automatic online platform to annotate A-to-I RNA editing sites in pre-mRNA splicing signals, microRNAs (miRNAs) and miRNA target untranslated regions (3' UTRs) from human (Homo sapiens) high-throughput sequencing data and predict their effects based on large-scale bioinformatic analysis. After analysing plenty of previously reported RNA editing events and human normal tissues RNA high-seq data, >60 000 potentially effective RNA editing events on functional genes were found. The RNA Editing Plus platform is available for free at https://www.rnaeditplus.org/, and we believe our platform governing multiple optimized methods will improve further studies of A-to-I-induced editing post-transcriptional regulation. Heming Wang, Yuanyuan Song, Tiedong Wang, Jimin Zhu, Xizhong Shen, Guangqi Song, Yicheng Zhao |
Briefings Bioinform. | 2 |
| 2017 | BioQueue: a novel pipeline framework to accelerate bioinformatics analysisabstractMOTIVATION: With the rapid development of Next-Generation Sequencing, a large amount of data is now available for bioinformatics research. Meanwhile, the presence of many pipeline frameworks makes it possible to analyse these data. However, these tools concentrate mainly on their syntax and design paradigms, and dispatch jobs based on users' experience about the resources needed by the execution of a certain step in a protocol. As a result, it is difficult for these tools to maximize the potential of computing resources, and avoid errors caused by overload, such as memory overflow. RESULTS: Here, we have developed BioQueue, a web-based framework that contains a checkpoint before each step to automatically estimate the system resources (CPU, memory and disk) needed by the step and then dispatch jobs accordingly. BioQueue possesses a shell command-like syntax instead of implementing a new script language, which means most biologists without computer programming background can access the efficient queue system with ease. AVAILABILITY AND IMPLEMENTATION: BioQueue is freely available at https://github.com/liyao001/BioQueue. The extensive documentation can be found at http://bioqueue.readthedocs.io. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Heming Wang, Yuanyuan Song, Guangchao Sui |
Bioinform. | 2 |
| 2014 | An improved differential box-counting method to estimate fractal dimensions of gray-level images
Heming Wang, Lanlan Jiang, Jiafei Zhao, Yuechao Zhao, Yongchen Song |
J. Vis. Commun. Image Represent. | 3 |