VLDB 2026 Research / reviewers in the wild / expert
Benlai Tang
dblp:198/9465
· DBLP profile ↗
7ranked-venue papers
0as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of Speech-Silence and Word-Punctuation
Jinzuomu Zhong, Korin Richmond, Zhiba Su, Benlai Tang, Fengjie Zhu |
INTERSPEECH | 8 |
| 2023 | CPNet: Exploiting CLIP-based Attention Condenser and Probability Map Guidance for High-fidelity Talking Face GenerationabstractRecently, talking face generation has drawn ever-increasing attention from the research community in computer vision due to its arduous challenges and widespread application scenarios, e.g. movie animation and virtual anchor. Although persevering efforts have been undertaken to enhance the fidelity and lip-sync quality of generated talking face videos, there is still large room for further improvements of synthesis quality and efficiency. Actually, these attempts somewhat ignore the explorations of fine-granularity feature extraction/integration and the consistency between probability distributions of landmarks, thereby recurring the issues of local details blurring and degraded fidelity. To mitigate these dilemmas, in this paper, a novel CLIP-based Attention and Probability Map Guided Network (CPNet) is delicately designed for inferring high-fidelity talking face videos. Specifically, considering the demands of fine-grained feature recalibration, a clip-based attention condenser is exploited to transfer knowledge with rich semantic priors from the prevailing CLIP model. Moreover, to guarantee the consistency in probability space and suppress the landmark ambiguity, we creatively propose the density map of facial landmark as auxiliary supervisory signal to guide the landmark distribution learning of generated frame. Extensive experiments on the widely-used benchmark dataset demonstrate the superiority of our CPNet against state of the arts in terms of image and lip-sync quality. In addition, a cohort of studies are also conducted to ablate the impacts of the individual pivotal components. Jingning Xu, Benlai Tang, Meirong Ma |
ICME | 2 |
| 2022 | Towards Using Clothes Style Transfer for Scenario-Aware Person Video GenerationabstractClothes style transfer for person video generation is a challenging task, due to drastic variations of intra-person appearance and video scenarios. To tackle this problem, most recent AdaIN-based architectures are proposed to extract clothes and scenario features for generation. However, these approaches suffer from being short of fine-grained details and are prone to distort the origin person. To further improve the generation performance, we propose a novel framework with disentangled multi-branch encoders and a shared decoder. Moreover, to pursue the strong video spatio-temporal consistency, an inner-frame discriminator is delicately designed with input being cross-frame difference. Besides, the proposed frame-work possesses the property of scenario adaptation. Extensive experiments on the TEDXPeople benchmark demonstrate the superiority of our method over state-of-the-art approaches in terms of image quality and video coherence. Jingning Xu, Benlai Tang, Siyuan Bian, Wenyi Guo, Xiang Yin 0006, Zejun Ma 0001 |
ICASSP | 2 |
| 2022 | Towards high-fidelity singing voice conversion with acoustic reference and contrastive predictive codingabstractRecently, phonetic posteriorgrams (PPGs) based methods have been quite popular in non-parallel singing voice conversion systems. However, due to the lack of acoustic information in PPGs, style and naturalness of the converted singing voices are still limited. To solve these problems, in this paper, we utilize an acoustic reference encoder to implicitly model singing characteristics. We experiment with different auxiliary features, including mel spectrograms, HuBERT, and the middle hidden feature (PPG-Mid) of pretrained automatic speech recognition (ASR) model, as the input of the reference encoder, and finally find the HuBERT feature is the best choice. In addition, we use contrastive predictive coding (CPC) module to further smooth the voices by predicting future observations in latent space. Experiments show that, compared with the baseline models, our proposed model can significantly improve the naturalness of converted singing voices and the similarity with the target singer. Moreover, our proposed model can also make the speakers with just speech data sing. Benlai Tang, Xiang Yin 0006, Yuan Wan, Yibiao Yu, Zejun Ma 0001 |
INTERSPEECH | 3 |
| 2021 | PPG-Based Singing Voice Conversion with Adversarial Representation LearningabstractSinging voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily convert songs while keeping their naturalness and intonation. We build an end-to-end architecture, taking phonetic posteriorgrams (PPGs) as inputs and generating mel spectrograms. Specifically, we implement two separate encoders: one encodes PPGs as content, and the other compresses mel spectrograms to supply acoustic and musical information. To improve the performance on timbre and melody, an adversarial singer confusion module and a mel-regressive representation learning module are designed for the model. Objective and subjective experiments are conducted on our private Chinese singing corpus. Comparing with the baselines, our methods can significantly improve the conversion performance in terms of naturalness, melody, and voice similarity. Moreover, our PPG-based method is proved to be robust for noisy sources. Benlai Tang, Xiang Yin 0006, Yuan Wan, Chen Shen 0011, Zejun Ma 0001 |
ICASSP | 2 |
| 2021 | Towards Realistic Visual Dubbing with Heterogeneous SourcesabstractThe task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data sources of videos and audios, thus causing the failure to leverage heterogeneous data sufficiently. In practice, it may be intractable to collect the perfect homologous data in some cases, for example, audio-corrupted or picture-blurry videos. To explore this kind of data and support high-fidelity few-shot visual dubbing, in this paper, we novelly propose a simple yet efficient two-stage framework with a higher flexibility of mining heterogeneous data. Specifically, our two-stage paradigm employs facial landmarks as intermediate prior of latent representations and disentangles the lip movements prediction from the core task of realistic talking head generation. By this means, our method makes it possible to independently utilize the training corpus for two-stage sub-networks using more available heterogeneous data easily acquired. Besides, thanks to the disentanglement, our framework allows a further fine-tuning for a given talking head, thereby leading to better speaker-identity preserving in the final synthesized results. Moreover, the proposed method can also transfer appearance features from others to the target speaker. Extensive experimental results demonstrate the superiority of our proposed method in generating highly realistic videos synchronized with the speech over the state-of-the-art. Tianyi Xie, Liucheng Liao, Benlai Tang, Xiang Yin 0006, Jianfei Yang 0001, Jiali Yao, Yang Zhang 0088, Zejun Ma 0001 |
ACM Multimedia | 4 |
| 2016 | Application of pronunciation knowledge on phoneme recognition by LSTM neural networkabstractWhen applied for phoneme recognition, the Connectionist Temporal Classification (CTC) objective function allows a neural network to be trained with the phoneme level transcriptions of training utterances. A limitation of the CTC is that it can not be applied directly for network training with large speech corpora, since those corpora usually only have word level transcriptions. This work extends the CTC such that a novel objective function can be evaluated even if only the word level transcriptions are available. Furthermore, various pronunciation knowledge is adopted to construct pronunciation networks which can model the pronunciations of connected speech more accurately. When combined with a bidirectional Long Short-term Memory (LSTM) network, the extended CTC achieves a phoneme error rate of 18.3% on the LibriSpeech corpus. When various pronunciation knowledge is applied, the error rate is further reduced by 18.6% relatively. Yuqin Gan, Benlai Tang |
ICPR | 4 |