Hualei Wang

dblp:41/10607 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Speech recognition and synthesis · 53% Language models and text generation · 20% Vision and language · 9%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 100%
Computer networks
1 paper
Physical-layer communications · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
audio-language model
1.012026
Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning · AAAI 2026
Natural language and speech › Question answering and dialogue systems › multimodal question answering
audio question answering
1.012026
Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning · AAAI 2026
Natural language and speech › Speech recognition and synthesis › audio-language model
audio understanding
1.012026
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models · AAAI 2026
Computer vision › Vision and language › image captioning
dense captioning
1.012026
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
1.012026
Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory · ACL (1) 2026
Natural language and speech › Speech recognition and synthesis › audio-language model
large audio language models
1.012026
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models · AAAI 2026
Natural language and speech › Language models and text generation › LLM agents
long-term memory
1.012026
Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory · ACL (1) 2026
Natural language and speech › Speech recognition and synthesis › speech evaluation
speech quality assessment
1.012026
Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator · AAAI 2026
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech
1.012026
Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator · AAAI 2026
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
zero-shot text-to-speech
1.012026
Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator · AAAI 2026
Health and well-being technologies
sleep monitoring
0.912025
SleepSMC: Ubiquitous Sleep Staging via Supervised Multimodal Coordination · ICLR 2025
Health and well-being technologies › sleep monitoring
sleep staging
0.912025
SleepSMC: Ubiquitous Sleep Staging via Supervised Multimodal Coordination · ICLR 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.312026
Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator · AAAI 2026
Audio and music processing › audio-language processing
audio-language understanding
0.312026
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models · AAAI 2026
Physical-layer communications › signal processing for communications
transceiver design
0.212016
Transceiver designs with matrix-version water-filling architecture under mixed power constraints · Sci. China Inf. Sci. 2016
Physical-layer communications › power allocation
water-filling
0.212016
Transceiver designs with matrix-version water-filling architecture under mixed power constraints · Sci. China Inf. Sci. 2016
Physical-layer communications
power allocation
0.112016
Transceiver designs with matrix-version water-filling architecture under mixed power constraints · Sci. China Inf. Sci. 2016

Methods — techniques the papers use, named apart from their topics

temporal marker encoding · 2.0segment-level token merging · 2.0absolute time-aware encoding · 2.0speech correction · 1.0reward model · 1.0reinforcement learning · 1.0preference optimization · 1.0masking and regeneration · 1.0hierarchical graph · 1.0adaptive think accuracy reward · 1.0uncertainty estimation · 0.9multimodal learning · 0.9contrastive learning · 0.9matrix-version water-filling · 0.2
YearPublicationVenuePosition
2026 Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models
abstract
Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perception, leading to constrained capabilities on fine-grained tasks. We identify three key aspects that limit their temporal localization and long audio understanding: (i) timestamp representation, (ii) architecture, and (iii) data. To address this, we introduce TimeAudio, a novel method that empowers LALMs to connect their understanding of audio content with precise temporal perception. Specifically, we incorporate unique temporal markers to improve time-sensitive reasoning and apply an absolute time-aware encoding that explicitly grounds the acoustic features with absolute time information. Moreover, to realize end-to-end long audio understanding, we introduce a segment-level token merging module to substantially reduce audio token redundancy and enhance the efficiency of information extraction. Due to the lack of suitable datasets and evaluation metrics, we consolidate existing audio datasets into a new dataset focused on temporal tasks and establish a series of metrics to evaluate the fine-grained performance. Evaluations show strong performance across a variety of fine-grained tasks, such as dense captioning, temporal grounding, and timeline speech summarization, which demonstrates TimeAudio's robust temporal localization and reasoning capabilities.
Hualei Wang, Hong Liu 0007
AAAI1
2026 Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator
abstract
Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges, manifesting as mispronunciations, audible noise, and quality degradation. To address these issues, we introduce Vox-Evaluator, a multi-level evaluator designed to guide the correction of erroneous speech segments and preference alignment for TTS systems. It is capable of identifying the temporal boundaries of erroneous segments and providing a holistic quality assessment of the generated speech. Specifically, to refine erroneous segments and enhance the robustness of the zero-shot TTS model, we propose to automatically identify acoustic errors with the evaluator, mask the erroneous segments, and finally regenerate speech conditioning on the correct portions. In addition, the fine-gained information obtained from Vox-Evaluator can guide the preference alignment for TTS model, thereby reducing the bad cases in speech synthesize. Due to the lack of suitable training datasets for the Vox-Evaluator, we also constructed a synthesized text-speech dataset annotated with fine-grained pronunciation errors or audio quality issues. The experimental results demonstrate the effectiveness of the proposed Vox-Evaluator in enhancing the stability and fidelity of TTS systems through the speech correction mechanism and preference optimization.
Hualei Wang, Na Li 0012, Chuke Wang, Zhifeng Li 0001, Dong Yu 0001
AAAI1
2026 Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning
abstract
Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded substantial benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of achieving human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs through improved adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that assist the model in distinguishing between valid and flawed reasoning paths during training. Experimental results demonstrate that Audio-Thinker models outperform existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities.
Chenxing Li, Wenfu Wang, Hao Zhang 0112, Hualei Wang, Meng Yu 0003, Dong Yu 0001
AAAI5
2026 Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory
abstract
Zihao Tang, Xin Yu, Ziyu Xiao, Zengxuan Wen, Zelin Li, Jiaxi Zhou, Hualei Wang, Haohua Wang, Haizhen Huang, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ziyu Xiao, Zengxuan Wen, Hualei Wang, Haizhen Huang, Feng Sun 0008, Qi Zhang 0066
ACL (1)7
2025 SleepSMC: Ubiquitous Sleep Staging via Supervised Multimodal Coordination
abstract
Sleep staging is critical for assessing sleep quality and tracking health. Polysomnography (PSG) provides comprehensive multimodal sleep-related information, but its complexity and impracticality limit its practical use in daily and ubiquitous monitoring. Conversely, unimodal devices offer more convenience but less accuracy. Existing multimodal learning paradigms typically assume that the data types remain consistent between the training and testing phases. This makes it challenging to leverage information from other modalities in ubiquitous scenarios (e.g., at home) where only one modality is available. To address this issue, we introduce a novel framework for ubiquitous Sleep staging via Supervised Multimodal Coordination, called SleepSMC. To capture category-related consistency and complementarity across modality-level instances, we propose supervised modality-level instance contrastive coordination. Specifically, modality-level instances within the same category are considered positive pairs, while those from different categories are considered negative pairs. To explore the varying reliability of auxiliary modalities, we calculate uncertainty estimates based on the variance in confidence scores for correct predictions during multiple rounds of random masks. These uncertainty estimates are employed to assign adaptive weights to multiple auxiliary modalities during contrastive learning, ensuring that the primary modality learns from high-quality, category-related features. Experimental results on four public datasets, ISRUC-S3, MASS-SS3, Sleep-EDF-78, and ISRUC-S1, show that SleepSMC achieves state-of-the-art cross-subject performance. SleepSMC significantly improves performance when only one modality is present during testing, making it suitable for ubiquitous sleep monitoring.
Shuo Ma 0001, Yingwei Zhang 0002, Yiqiang Chen 0001, Hualei Wang, Wei Zhang 0082, Ziyu Jia
ICLR4
2025 Diverse Audio Caption Generation with Semantic-aware Diffusion Model
abstract
Audio captioning aims to perceive sound events in different ways and describe an audio clip from various perspectives. Most existing audio captioning methods tend to generate captions that are deterministic and simple, lacking diversity and limiting their applicability in real-world scenarios. Recently, diffusion-based methods have achieved significant progress in producing diverse captions, but with a potential compromise in accuracy and fluency due to the abstract nature of language and the variable supervised targets. In this work, we propose a semantic-aware diffusion model that leverages its intrinsic stochastic sampling and global context to generate captions. In order to maintain accuracy, we integrate the global CLAP embeddings into the denoising process of the diffusion model to serve as semantic context. To further enhance diversity, we propose an optimized inference process that incorporates a dynamic denoising strategy during the token-restored stage. Extensive experiments on the AudioCaps and Clotho dataset demonstrate that our model achieves superior results on accuracy and diversity metrics compared to state-of-the-art diverse audio caption methods.
Hualei Wang, Hong Liu 0007
ICME1
2024 Leveraging Language Model Capabilities for Sound Event Detection
Hualei Wang, Jianguo Mao, Zhifang Guo, Jiarui Wan, Hong Liu 0007
INTERSPEECH1
2024 SF-SNF: A Small-file Sniffer System for Hadoop Clusters
abstract
Apache Hadoop has been a major component in the big data ecosystem for more than a decade. It relies on the Hadoop Distributed File System (HDFS) to store large datasets and MapReduce to process these distributed datasets. HDFS manages the metadata of all its files through a server known as the Namenode. To achieve high availability (HA), Hadoop clusters typically deploy two Namenodes: one active and one standby. This architecture enables Hadoop to store and process massive files with good reliability. However, HDFS often encounters significant performance degradation when managing a large number of small files. Scanning all files to locate the small ones through the Namenode’s service becomes time-consuming and adds extra burdens to the Namenodes. There is a lack of research on how to identify the hotspots of small files in HDFS without querying the Namenode in a time-efficient manner. In this paper, we designed a big data system to identify small files in HDFS by parsing the File System Image (FSImage), which is generated periodically on the Namenode. This system utilizes the standby Namenode to parse the FSImage and send the file information to an Apache Kafka topic. A Doris Routine Load procedure then listens to the topic and loads the information into a table containing the file information for real-time querying by users. This approach allows Hadoop cluster administrators to identify small files using the standby Namenode without impacting the normal operation of the Hadoop cluster. Additionally, it can serve as a tool to locate small-file hotspot tables in a Hive data warehouse.
Hualei Wang
ISPA3
2016 Transceiver designs with matrix-version water-filling architecture under mixed power constraints
Chengwen Xing, Zesong Fei, Yiqing Zhou 0001, Zhengang Pan, Hualei Wang
Sci. China Inf. Sci.5
2015 Distributed optimization for downlink broadband small cell networks
abstract
Small cell networks have been recognized as a promising technology to realize high spectrum and energy efficiency communications in future wireless networks. However, the capacity of small cell networks is limited by the interference among the links communicating simultaneously. Efficient resource allocation and effective interference management is definitely imperative for small cells. In this paper, we propose an algorithm to optimize the power and subcarrier allocation jointly in order to maximize the weighted sum rate for dense small cell networks. Facing with a large amount of small cells, the optimization problem is in nature a large scale optimization problem. Using advanced decomposition theory, the proposed algorithm can effectively decompose the considered optimization problem into a series of much simpler subproblems which can be efficiently solved in parallel. Finally, simulation results demonstrate that the proposed algorithm enjoys greater performance gain and faster convergence as compared to the existing schemes.
Shaozhen Guo, Chengwen Xing, Zesong Fei, Hualei Wang, Zhengang Pan
ICC4
2014 A Temporal Domain Based Method against Pilot Contamination for Multi-Cell Massive MIMO Systems
abstract
Massive multiple-input and multiple-output (MIMO) systems have draw a lot of attention, due to its outstanding performance in energy saving and spectrum efficiency [1]. In massive MIMO systems, the intra-cell interference, uncorrelated noise and fast fading nearly vanish. Further study reveals that for multi-cell massive MIMO systems the performance will be limited by so-called pilot contamination (inter-cell interference) [2], where the transmitted CSI at one BS includes the channel coefficients from all the users using the same pilot. However, the analysis in [2] ignores a fact that the spatial/temporal characters of channel coefficients of different users are different such that they are distinguishable in certain extends. In practice, utterly isolating the user channels is extremely hard or even not feasible in certain cases. This paper firstly presents a practical temporal domain based method against pilot contamination without coordination, which just remains the strongest channel impulse responses (CIR) as the effective channel information. Simulation results show the outperformance of the proposed method with low complexity.
Hualei Wang, Zhengang Pan, Jiqing Ni, Sen Wang 0005, Chih-Lin I
VTC Spring1
2012 Unitary precoder design for MIMO spatial multiplexing systems with limited feedback
abstract
This paper studies limited feedback unitary precoding for MIMO spatial multiplexing systems with QR-SIC receiver and equal power allocation. An optimal codebook design criterion, which can design the optimal codebooks for various fading and spatial correlated channels with arbitrary antenna configurations, is proposed. However, due to the fact that the closed-form of the probability density function (p.d.f.) of the optimal precoding matrix is hard to obtain, the optimal codebook design criterion is difficult to be applied to the practical systems. Thus, a suboptimal codebook design criterion is given, which can be implemented in practice and is also applicable to any scenario. Sequentially, an iterative codebook design algorithm that converges to an optimum codebook is given. Simulation results demonstrate that the proposed scheme achieves better performance than the 3GPP R10 LTE-A codebooks.
Hualei Wang, Lihua Li 0001, Markku Juntti, Teemu Puotinen
CCNC1
2011 Multi-Cell Collaborative Transmission Combining Closed-Loop and Open-Loop Techniques
abstract
In this paper, we evaluate the existing transmission schemes in multi-cell environment, such as single-cell transmission with Inter-cell interference (ICI) as well as Collaborative Multi-Point transmission/reception (CoMP) technique in LTE-A, and propose an more effective solution to deal with the inter-cell interference. The proposed transmission structure takes advantage of the cooperation between base stations (BS), and combines both closed-loop and open-loop techniques. With the concept of Effective Channel, the classic Space Frequency Block Coding (SFBC) decode scheme can be used directly in our receiver for single layer transmission. Thus, we prove that the scenario of multiple base stations can employ SFBC conveniently while the performance is guaranteed. Analysis and simulation results show that our method outperforms conventional single cell transmission. Moreover, the proposed transmission scheme can obtain elegant performance enhancement compared with local precoding and global precoding of CoMP in terms of capacity, and doubly reduce the feedback overhead at the same time.
Ji Wang 0004, Lihua Li 0001, Hualei Wang, Qi Sun 0001, Wanlu Sun
VTC Fall4
2011 A Transmit Precoding Scheme for Downlink Multiuser MIMO Systems
abstract
This paper focuses on transmit precoding for multiuser MIMO downlink systems. In multiuser MIMO systems, multiuser interference (MUI) and noise are two well-known factors with respect to the system's performance. Block diagonalization (BD) method is proposed to completely eliminate MUI by placing the intended users in the nullspace of all the unintended users. But the BD method imposes a condition on the relation between the number of transmit and receive antennas. In addition, BD method causes the noise enhancement due to not considering noise's influence. Thus, at low and medium signal-to-noise ratio (SNR) regime, the performance of the BD scheme is poor. In this paper, we propose a novel precoding approach for users with multiple antennas to overcome the above mentioned drawbacks of the BD method for multiuser MIMO precoding systems. Simulation results confirm that the proposed algorithm achieves performance improvement over the conventional BD scheme with low complexity. The effect of channel estimate errors on system performance is also studied.
Hualei Wang, Lihua Li 0001, Ji Wang 0004
VTC Fall1