VLDB 2026 Research / reviewers in the wild / expert
Daipeng Zhang 0001
dblp:214/5747-1
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0006-8199-2217ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 64% Trustworthy machine learning · 28% Language models and text generation · 8% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
hallucination |
1.0 | 1 | 2026 | HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026 |
Natural language and speech › Speech recognition and synthesis
speech reconstruction |
1.0 | 1 | 2026 | EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment · AAAI 2026 |
Audio and music processing › audio analysis
audio understanding |
1.0 | 1 | 2026 | HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026 |
Natural language and speech › Speech recognition and synthesis
audio-language model |
0.3 | 1 | 2026 | HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2026 | HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
benchmark evaluation · 2.0adversarial prompting · 2.0variational autoencoder · 1.0encoding alignment · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding AlignmentabstractDysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparities in the feature space between normal and dysarthric speech may result in suboptimal speech reconstruction, thereby degrading speech intelligibility. To enhance the reconstruction ability of speech feature spaces, this paper proposes a DSR model named the Encoding-Aligned Variational Autoencoder (EA-VAE). By incorporating alignment modules of frame-level embedding features, prior distributions, and duration into the encoder of the VAE, the model explicitly aligns the dysarthric speech encoding with a representation of the parallel normal speech. A shared decoder is then used to generate speech with improved intelligibility. Experimental results on the UASpeech benchmark confirm that EA-VAE achieves state-of-the-art performance, with a 31.7% relative word error rate reduction and the highest subjective MOS score (4.48), thoroughly validating the effectiveness and advancements of the proposed method in dysarthric speech reconstruction. Daipeng Zhang 0001, Wenhuan Lu, Xianghu Yue, Hongcheng Zhang, Jianguo Wei |
AAAI | 1 |
| 2026 | HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language ModelsabstractLarge Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks.However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain.Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music.HalluAudio comprises over 5K humanverified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA.To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions.Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes.We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music.Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs. Feiyu Zhao, Wenhuan Lu, Daipeng Zhang 0001, Xianghu Yue, Jianguo Wei |
ACL (1) | 4 |