Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wenfu Wang

dblp:180/2589 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-1254-3910ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 33% Question answering and dialogue systems · 33% Language models and text generation · 33%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
audio-language model
1.012026
Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning · AAAI 2026
Natural language and speech › Question answering and dialogue systems › multimodal question answering
audio question answering
1.012026
Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning · AAAI 2026

Methods — techniques the papers use, named apart from their topics

reward model · 1.0reinforcement learning · 1.0adaptive think accuracy reward · 1.0
YearPublicationVenuePosition
2026 Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning
abstract
Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded substantial benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of achieving human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs through improved adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that assist the model in distinguishing between valid and flawed reasoning paths during training. Experimental results demonstrate that Audio-Thinker models outperform existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities.
Chenxing Li, Wenfu Wang, Hao Zhang 0112, Hualei Wang, Meng Yu 0003, Dong Yu 0001
AAAI3
2025 Toward Camera Open-Set 3D Object Detection for Autonomous Driving Scenarios
abstract
Conventional camera-based 3D object detectors in autonomous driving are limited to recognizing a predefined set of objects, which poses a safety risk when encountering novel or unseen objects in real-world scenarios. To address this limitation, we present OS-Det3D, a two-stage training framework designed for camera-based open-set 3D object detection. In the first stage, our proposed 3D object discovery network (ODN3D) uses geometric cues from LiDAR point clouds to generate class-agnostic 3D object proposals, each of which are assigned a 3D objectness score. This approach allows the network to discover objects beyond known categories, allowing for the detection of unfamiliar objects. However, due to the absence of class constraints, ODN3D-generated proposals may include noisy data, particularly in cluttered or dynamic scenes. To mitigate this issue, we introduce a joint selection (JS) module in the second stage. The JS module uses both camera bird’s eye view (BEV) feature responses and 3D objectness scores to filter out low-quality proposals, yielding high-quality pseudo ground truth for unknown objects. OS-Det3D significantly enhances the ability of camera 3D detectors to discover and identify unknown objects while also improving the performance on known objects, as demonstrated through extensive experiments on the nuScenes and KITTI datasets.
Zhuolin He, Xinrun Li, Jiacheng Tang, Shoumeng Qiu, Wenfu Wang, Xiangyang Xue 0001, Jian Pu
IEEE Trans. Intell. Transp. Syst.5
2023 Elastic Knowledge Distillation by Learning From Recollection
abstract
Model performance can be further improved with the extra guidance apart from the one-hot ground truth. To achieve it, recently proposed recollection-based methods utilize the valuable information contained in the past training history and derive a "recollection" from it to provide data-driven prior to guide the training. In this article, we focus on two fundamental aspects of this method, i.e., recollection construction and recollection utilization. Specifically, to meet the various demands of models with different capacities and at different training periods, we propose to construct a set of recollections with diverse distributions from the same training history. After that, all the recollections collaborate together to provide guidance, which is adaptive to different model capacities, as well as different training periods, according to our similarity-based elastic knowledge distillation (KD) algorithm. Without any external prior to guide the training, our method achieves a significant performance gain and outperforms the methods of the same category, even as well as KD with well-trained teacher. Extensive experiments and further analysis are conducted to demonstrate the effectiveness of our method.
Yongjian Fu 0002, Hanbin Zhao, Wenfu Wang, Weihao Fang, Yueting Zhuang, Xi Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 SDAPNet: End-to-End Multi-task Simultaneous Detection and Prediction Network
abstract
Accurate detection of objects in perception system is a basic task for autonomous vehicles to operate reliably. The current mainstream paradigms for perception and prediction are able to decompose into three sequential subtasks: detection, tracking and prediction. Much of the previous works focus on one of these subtasks and internal relationship between the various subtasks are therefore ignored. In this work we propose an end-to-end model that is able to jointly carry out 3D detection and motion prediction in the context of autonomous vehicles, which can benefit from joint optimization of two tasks. Considering computational consumption and memory cost, our approach utilizes 2D convolutions across space over a bird's eye view (BEV) representation of the 3D world, not 3D convolutions which require more inference time for a single point cloud. In this paper, the key model is the sequential multi-frame point clouds channels based on BEV, which can help perform object detection and learn temporal information at the same time. Experiments on nuScenes dataset show that the proposed network achieves state-of-the-art results, and the trade-off between accuracy and efficiency of two tasks are obtained. Additionally, we validate the effectiveness of our models through the ablation studies. Specifically, by sharing network, parameters and detection results we can perform two tasks within 32 FPS.
Shanding Ye, Wenfu Wang, Yongjian Fu 0002
IJCNN3
2017 Combining unidirectional long short-term memory with convolutional output layer for high-performance speech synthesis
abstract
In this paper, we target improving the accuracy of acoustic modelling for statistical parametric speech synthesis (SPSS) and introduce the convolutional neural network (CNN) due to its powerful capacity in locality modelling. A novel model architecture combining unidirectional long short-term memory (LSTM) and a time-domain convolutional output layer (COL) is proposed and employed to acoustic modelling. The two components complement each other and result in a high-performance synthesis system. Specifically, the unidirectional LSTM can learn expressive feature representations from history context and the COL ingeniously absorbs some of these representations within a look-ahead window to advance predictions. This complementary mechanism significantly improve the predictive accuracy and the quality of synthetic speech. In addition, the unique operation mechanism of convolution makes COL a fine parameter trajectory smoother between consecutive frames. Subjective preference tests show that the proposed architecture can synthesize natural sounding speech without dynamic features.
Wenfu Wang, Bo Xu 0002
ICASSP1
2016 Gating recurrent mixture density networks for acoustic modeling in statistical parametric speech synthesis
abstract
Though recurrent neural networks (RNNs) using long short-term memory (LSTM) units can address the issue of long-span dependencies across the linguistic inputs and have achieved the state-of-the-art performance for statistical parametric speech synthesis (SPSS), another limitation of the intrinsic uni-Gaussian nature of mean square error (MSE) objective function still remains. This paper proposes a gating recurrent mixture density network (GRMDN) architecture to jointly address these two problems in neural network based SPSS. What's more, the gated recurrent unit (GRU), which is much simpler and has more intelligible work mechanism than LSTM, is also investigated as an alternative gating unit in RNN based acoustic modeling. Experimental results show that the proposed GRMDN architecture can synthesize more natural speech than its MSE-trained counterpart and both the two gating units (LSTM and GRU) show comparable performance.
Wenfu Wang, Bo Xu 0002
ICASSP1
2016 End-to-End Language Identification Using Attention-Based Recurrent Neural Networks
Wang Geng, Wenfu Wang, Xinyuan Cai, Bo Xu 0002
INTERSPEECH2
2016 Gating Recurrent Enhanced Memory Neural Networks on Language Identification
Wang Geng, Wenfu Wang, Xinyuan Cai, Bo Xu 0002
INTERSPEECH3
2016 First Step Towards End-to-End Parametric TTS Synthesis: Generating Spectral Parameters with Neural Attention
Wenfu Wang, Bo Xu 0002
INTERSPEECH1