Daipeng Zhang 0001

dblp:214/5747-1 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0006-8199-2217ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Speech recognition and synthesis · 64% Trustworthy machine learning · 28% Language models and text generation · 8%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
hallucination
1.012026
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026
Natural language and speech › Speech recognition and synthesis
speech reconstruction
1.012026
EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment · AAAI 2026
Audio and music processing › audio analysis
audio understanding
1.012026
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026
Natural language and speech › Speech recognition and synthesis
audio-language model
0.312026
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model
0.312026
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

benchmark evaluation · 2.0adversarial prompting · 2.0variational autoencoder · 1.0encoding alignment · 1.0
YearPublicationVenuePosition
2026 EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment
abstract
Dysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparities in the feature space between normal and dysarthric speech may result in suboptimal speech reconstruction, thereby degrading speech intelligibility. To enhance the reconstruction ability of speech feature spaces, this paper proposes a DSR model named the Encoding-Aligned Variational Autoencoder (EA-VAE). By incorporating alignment modules of frame-level embedding features, prior distributions, and duration into the encoder of the VAE, the model explicitly aligns the dysarthric speech encoding with a representation of the parallel normal speech. A shared decoder is then used to generate speech with improved intelligibility. Experimental results on the UASpeech benchmark confirm that EA-VAE achieves state-of-the-art performance, with a 31.7% relative word error rate reduction and the highest subjective MOS score (4.48), thoroughly validating the effectiveness and advancements of the proposed method in dysarthric speech reconstruction.
Daipeng Zhang 0001, Wenhuan Lu, Xianghu Yue, Hongcheng Zhang, Jianguo Wei
AAAI1
2026 HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
abstract
Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks.However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain.Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music.HalluAudio comprises over 5K humanverified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA.To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions.Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes.We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music.Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs.
Feiyu Zhao, Wenhuan Lu, Daipeng Zhang 0001, Xianghu Yue, Jianguo Wei
ACL (1)4