EDBT 2026 Demo / reviewers in the wild / expert
Fengping Wang
dblp:179/8081
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource ScenariosabstractZero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal content separately, losing essential acoustic information that degrades output quality while requiring significant computational resources. To overcome these limitations, we propose HQ-SVC, an efficient framework for high-quality zero-shot SVC. HQ-SVC first extracts jointly content and speaker features using a decoupled codec. It then enhances fidelity through pitch and volume modeling, preserving critical acoustic information typically lost in separate modeling approaches, and progressively refines outputs via differentiable signal processing and diffusion techniques. Evaluations confirm HQ-SVC significantly outperforms state-of-the-art zero-shot SVC methods in conversion quality and efficiency. Beyond voice conversion, HQ-SVC achieves superior voice naturalness compared to specialized audio super-resolution methods while natively supporting voice super-resolution tasks. Bingsong Bai, Yizhong Geng, Fengping Wang, Puyuan Guo, Yingming Gao, Ya Li 0001 |
AAAI | 3 |
| 2026 | Robust traffic sign detection in real-world harsh conditions: A pioneering benchmark dataset and attention-based methodology
Fengping Wang, Meng Wang 0015, Baobao Liu, Haiwei Xue |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | A micro-expression recognition algorithm fusing visual information with textual semantics
Fengping Wang, Jie Li 0089, Chun Qi, Pan Wang 0004 |
Expert Syst. Appl. | 1 |
| 2025 | DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speechabstractTraditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often introduce data inconsistencies and increase complexity. To address these issues, we propose DetailTTS, a zero-shot TTS system based on a conditional variational autoencoder. It incorporates two key components: the Prior Detail Module and the Duration Detail Module, which capture residual detail information missed during alignment. These modules effectively enhance the model’s ability to retain fine-grained details, significantly improving speech quality while simplifying the model by obviating the need for additional inputs. Experiments on the WenetSpeech4TTS dataset show that DetailTTS outperforms traditional TTS systems in both naturalness and speaker similarity, even in zero-shot scenarios. Our source code and demo page are available at https://detailtts.github.io/. Yichen Han, Yizhong Geng, Yingming Gao, Fengping Wang, Bingsong Bai, Jinlong Xue, Yayue Deng, Zhengqi Wen, Ya Li 0001 |
ICASSP | 5 |
| 2025 | A neighbor-aware feature enhancement network for crowd counting
Jie Li 0089, Chun Qi, Runrun Zou, Fengping Wang, Pan Wang 0004 |
Image Vis. Comput. | 6 |
| 2025 | A multi-modal multi-scale network based on Transformer for micro-expression recognition
Fengping Wang, Jie Li 0089, Chun Qi, Pan Wang 0004 |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | Progressive Crowd Enhancement De-Background Network for crowd counting
Jie Li 0089, Chun Qi, Fengping Wang, Pan Wang 0004 |
Vis. Comput. | 4 |
| 2024 | Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech SynthesisabstractConversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While previous methods have already delved into enhancing context comprehension, context representation still lacks effective representation capabilities and context-sensitive discriminability. In this paper, we introduce a contrastive learning-based CSS framework, CONCSS. Within this framework, we define an innovative pretext task specific to CSS that enables the model to perform self-supervised learning on unlabeled conversational datasets to boost the model’s context understanding. Additionally, we introduce a sampling strategy for negative sample augmentation to enhance context vectors’ discriminability. This is the first attempt to integrate contrastive learning into CSS. We conduct ablation studies on different contrastive learning strategies and comprehensive experiments in comparison with prior CSS systems. Results demonstrate that the synthesized speech from our proposed method exhibits more contextually appropriate and sensitive prosody. Yayue Deng, Jinlong Xue, Yukang Jia, Yichen Han, Fengping Wang, Yingming Gao, Dengfeng Ke, Ya Li 0001 |
ICASSP | 6 |
| 2024 | SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
Bingsong Bai, Fengping Wang, Yingming Gao, Ya Li 0001 |
INTERSPEECH | 2 |
| 2024 | A 4D spontaneous micro-expression database: Establishment and evaluationabstractAbstract Micro‐expressions are spontaneous and unconscious facial movements that reveal individuals’ genuine inner emotions. They hold significant potential in various psychological testing fields. As the face is a 3D deformation object, the emergence of facial expression leads to spatial deformation of the face. However, existing databases primarily offer 2D video sequences, limiting descriptions of 3D spatial information related to micro‐expressions. Here, a new micro‐expression database is proposed, which contains 2D image sequences and corresponding 3D point cloud sequences. These samples were classified using both an objective method based on the facial action coding system and a non‐objective emotion classification method that considers video contents and participants’ self‐reports. A variety of feature extraction techniques are applied to 2D data, including traditional algorithms and deep learning methods. Additionally, a novel local curvature‐based algorithm is developed to extract 3D spatio‐temporal deformation features from the 3D data. The authors evaluated the classification accuracies of these two features individually and their fusion results under leave‐one‐subject‐out (LOSO) and tenfold cross‐validation. The results demonstrate that fusing 3D features with 2D features results in improved recognition performance compared to using 2D features alone. Fengping Wang, Jie Li 0089, Chun Qi, Pan Wang 0004 |
IET Image Process. | 1 |
| 2024 | JGULF: Joint global and unilateral local feature network for micro-expression recognition
Fengping Wang, Jie Li 0089, Chun Qi, Pan Wang 0004 |
Image Vis. Comput. | 1 |
| 2024 | YOLO-SG: Small traffic signs detection method in complex scene
Yanjiang Han, Fengping Wang, Jianyang Zhang |
J. Supercomput. | 2 |
| 2024 | DM-YOLOX aerial object detection method with intensive attention mechanism
Fengping Wang, Yanjiang Han, Jianyang Zhang |
J. Supercomput. | 2 |
| 2023 | M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech SynthesisabstractConversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational TTS systems only focus on extracting global information and omit local prosody features, which contain important fine-grained information like keywords and emphasis. Moreover, it is insufficient to only consider the textual features, and acoustic features also contain various prosody information. Hence, we propose M2-CTTS, an end-to-end multi-scale multi-modal conversational text-to-speech system, aiming to comprehensively utilize historical conversation and enhance prosodic expression. More specifically, we design a textual context module and an acoustic context module with both coarse-grained and fine-grained modeling. Experimental results demonstrate that our model mixed with fine-grained context information and additionally considering acoustic features achieves better prosody performance and naturalness in CMOS tests. Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li 0001, Yingming Gao, Jianhua Tao 0001, Jianqing Sun, Jiaen Liang |
ICASSP | 3 |
| 2023 | CMCU-CSS: Enhancing Naturalness via Commonsense-based Multi-modal Context Understanding in Conversational Speech SynthesisabstractConversational Speech Synthesis (CSS) aims to produce speech appropriate for oral communication. However, the complexity of context dependency modeling poses significant challenges in the field of CSS, especially the mutual psychological influence between interlocutors. Previous studies have verified that prior commonsense knowledge helps machines understand subtle psychological information (e.g., feelings and intentions) in spontaneous oral dialogues. Therefore, to enhance context understanding and improve the naturalness of synthesized speech, we propose a novel conversational speech synthesis system (CMCU-CSS) that incorporates the Commonsense-based Multi-modal Context Understanding (CMCU) module to model the dynamic emotional interaction among interlocutors. Specifically, we first utilize three implicit states (intent state, internal state and external state) in CMCU to model the context dependency between inter/intra speakers with the help of commonsense knowledge. Furthermore, we infer emotion vectors from the fusion of these implicit states and multi-modal features to enhance the emotion discriminability of synthesized speech. This is the first attempt to combine commonsense knowledge with conversational speech synthesis, and its effect in terms of emotion discriminability of synthetic speech is evaluated by emotion recognition in conversation task. The results of subjective and objective evaluations demonstrate that the CMCU-CSS model achieves more natural speech with context-appropriate emotion and is equipped with the best emotion discriminability, surpassing that of other conversational speech synthesis models. Yayue Deng, Jinlong Xue, Fengping Wang, Yingming Gao, Ya Li 0001 |
ACM Multimedia | 3 |
| 2023 | Multi-Scale and spatial position-based channel attention network for crowd counting
Jie Li 0089, Chun Qi, Pan Wang 0004, Fengping Wang |
J. Vis. Commun. Image Represent. | 6 |
| 2020 | Mapping Road Based on Multiple Features and B-GVF SnakeabstractAs a significant application in aerial image, road mapping is still a difficult task since roads show complex features caused by the influence of spectral reflectance, shadows and occlusions. To achieve a satisfying result, a new method combing multiple road features and biased gradient vector flow (B-GVF) snake is studied in this paper. First, an exponential function is applied to fuse the color-based and structure-based measure for gaining the saliency maps which is viewed as the candidate region of B-GVF snake; Secondly, the initial road boundary is calculated from the candidate region using a region-growing algorithm, and then an gradient map is produced by an Gaussian filtering function; at last, a normally biased GVF external force is proposed for mapping road edges, which keeps the diffusion along the tangential direction of the isophotes and biases along the normal direction. Experimental results show that the proposed approach has the good performance in Completeness, Correctness, and F-measure comparing with other state-of-the-art methods. Fengping Wang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2019 | Road extraction using modified dark channel prior and neighborhood FCM in foggy aerial images
Fengping Wang |
Multim. Tools Appl. | 1 |