EDBT 2026 Demo / reviewers in the wild / expert
Pengfei Hu 0004
dblp:71/9969-4
· DBLP profile ↗
12ranked-venue papers
0as first author
11since 2021 · last 2024
0009-0000-4537-6288ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Common Sense Language-Guided Exploration and Hierarchical Dense Perception for Instruction Following Embodied AgentsabstractEmbodied Instruction Following (EIF) involves the task of locating and manipulating objects according to language instructions. Existing methods face challenges in small object navigation due to ineffective exploration and imperfect perception, which ultimately affects their performance. This study focuses on small object navigation in the EIF domain. We propose Common Sense Language-guided exploration (CSL), a novel approach that leverages common-sense knowledge from seen scenes and information from language instructions to infer the location of objects. The proposed CSL significantly improves exploration efficiency. Additionally, we propose Hierarchical Dense Perception (HDP), which uses hierarchical features to perform semantic segmentation and depth estimation. The use of HDP significantly improves the agent’s perceptual capabilities. Experiments on the ALFRED benchmark demonstrate the effectiveness of CSL-HDP. The proposed CSL-HDP achieves an absolute improvement of 9.29% (18.45% relative) on unseen test scenes compared to the previous state-of-the-art, securing the top position on the leaderboard. Code will be available at https://github.com/Cyuanwen/CSL-HDP. Yuanwen Chen, Yaran Chen, Dongbin Zhao, Yunzhen Zhao, Pengfei Hu 0004 |
ICME | 7 |
| 2024 | ATV3D: 3D Object Detection from Attention-based Three-view RepresentationabstractIn the fields of autonomous driving and robot perception, the majority of methods are designed for onboard camera object detection, while there are fewer methods specifically tailored to environmental cameras. However, environmental cameras have the capability to capture a significant amount of road geometry and vehicle position information, which can enhance the safety of autonomous driving. Nevertheless, there is a difference in perspective between environmental cameras and onboard cameras, resulting in poorer performance of many methods designed for onboard camera 3D object detection when apply to environmental camera. In this paper, we propose a 3D Object Detection Algorithm from Attention-based Three-view Representation (ATV3D). The algorithm projects the 2D image features onto three orthogonal views (left view, front view, bird’s eye view) to achieve a representation of the 3D information. Compared to voxel-based 3D detection methods, our proposed approach retains the ability to capture 3D features while reducing computational complexity. During the process of three-view representation, we design a feature projection module based on attention. Unlike inverse perspective mapping that requires precise camera parameters, the attention can implicitly learn the mapping relationship from 2D images to the three-view planes. This enables the extraction and transformation of image features without the calibrated camera parameters, effectively addressing challenges associated with obtaining camera parameters for environmental cameras and their susceptibility to natural factors. The experimental results on the DAIR-V2X dataset demonstrate that our method achieves a 3D detection mean average precision (mAP) of 73.6%, surpassing the performance of previous calibration-free environmental camera methods. Furthermore, our method achieves the highest detection accuracy on the indoor multi-view robot dataset Neurons Perception, providing evidence of its outstanding detection performance. Yaran Chen, Haoran Li 0010, Yunzhen Zhao, Pengfei Hu 0004 |
IJCNN | 6 |
| 2024 | Decoupling and Interacting Multi-Task Learning Network for Joint Speech and Accent RecognitionabstractAccents pose significant challenges for speech recognition systems. Although joint automatic speech recognition (ASR) and accent recognition (AR) training has been proven effective in handling multi-accent scenarios, current multi-task ASR-AR approaches overlook the granularity differences between tasks. Fine-grained units capture pronunciation-related accent characteristics, while coarse-grained units are better for learning linguistic information. Moreover, an explicit interaction of two tasks can provide complementary information and improve the other's performance, but it is rarely used by existing approaches. In this paper, we propose a novel Decoupling and Interacting Multi-task Network (DIMNet) for joint speech and accent recognition, which is comprised of a connectionist temporal classification (CTC) branch, an AR branch, an ASR branch, and a bottom feature encoder. Specifically, AR and ASR are first decoupled by separated branches and two-granular modeling units to learn task-specific representations. The AR branch is from our previously proposed linguistic-acoustic bimodal AR model and the ASR branch is an encoder-decoder based Conformer model. Then, for the task interaction, the CTC branch provides aligned text for the AR task, while accent embeddings extracted from our AR model are incorporated into the ASR branch's encoder and decoder. Finally, during ASR inference, a cross-granular rescoring method is introduced to fuse the complementary information from the CTC and attention decoder after the decoupling. Our experiments on English and Chinese datasets demonstrate the effectiveness of the DIMNet, which achieves${21.45\%}$/${28.53\%}$AR accuracy relative improvement and${32.33\%}$/${14.55\%}$ASR error rate relative reduction over a published standard baseline, respectively. Qijie Shao, Jinghao Yan, Pengfei Hu 0004, Lei Xie 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | A Method of Audio-Visual Person Verification by Mining Connections between Time Series
Peiwen Sun, Zishan Liu, Yougen Yuan, Taotao Zhang, Honggang Zhang 0002, Pengfei Hu 0004 |
INTERSPEECH | 7 |
| 2022 | Fake Audio Detection Based On Unsupervised Pretraining ModelsabstractThis work presents our systems for the ADD2022 challenge. The ADD2022 challenge is the first audio deep synthesis detection challenge, which aims to spot various kinds of fake audios. We have explored using unsupervised pretraining models to build fake audio detection systems. Results indicate that unsupervised pretraining models can achieve excellent performance for fake audio detection. Our final EER results for low-quality fake audio detection and partially fake audio detection are 32.80% and 4.80% relatively. For partially fake audio detection, our results ranked first in the competition. Even trained with totally mismatched data, our method still generalizes well for partially fake audio detection. Zhiqiang Lv, Pengfei Hu 0004 |
ICASSP | 4 |
| 2022 | ICPR 2022 Challenge on Multi-Modal Subtitle RecognitionabstractVideo subtitle recognition, as one of the basic elements of video editing, has received increasing attention recently. However, the misaligned between audio and subtitles as well as the costly manual annotation remain a demanding issue toward subsequent intelligent processing. In this paper, we introduce a multi-modal subtitle recognition challenge for ICPR 2022, in which we present a large-scale video dataset (215 hours in total for visual and audio annotations), and 3 tracks including: 1) extracting subtitles in visual modality with audio annotation (ESV); 2) extracting subtitles in audio modality with visual annotations (ESA); and 3) extracting subtitles with both visual and audio annotation (ESVA). The challenge attracts 376 participants, among which the methods of top 3 teams on each track have been elaborated. Shen Huang, Pengfei Hu 0004, Jian Kang 0006, Weida Liang, Yaqiang Wu, Yong Liu 0027 |
ICPR | 4 |
| 2022 | PM-MMUT: Boosted Phone-mask Data Augmentation using Multi-Modeling Unit Training for Phonetic-Reduction-Robust E2E Speech RecognitionabstractConsonant and vowel reduction are often encountered in speech, which might cause performance degradation in automatic speech recognition (ASR).Our recently proposed learning strategy based on masking, Phone Masking Training (PMT), alleviates the impact of such phenomenon in Uyghur ASR.Although PMT achieves remarkably improvements, there still exists room for further gains due to the granularity mismatch between the masking unit of PMT (phoneme) and the modeling unit (word-piece).To boost the performance of PMT, we propose multi-modeling unit training (MMUT) architecture fusion with PMT (PM-MMUT).The idea of MMUT framework is to split the Encoder into two parts including acoustic feature sequences to phoneme-level representation (AF-to-PLR) and phoneme-level representation to word-piece-level representation (PLR-to-WPLR).It allows AF-to-PLR to be optimized by an intermediate phoneme-based CTC loss to learn the rich phoneme-level context information brought by PMT.Experimental results on Uyghur ASR show that the proposed approaches outperform obviously the pure PMT.We also conduct experiments on the 960-hour Librispeech benchmark using ES-Pnet1, which achieves about 10% relative WER reduction on all the test set without LM fusion comparing with the latest official ESPnet1 pre-trained model. Pengfei Hu 0004, Nurmemet Yolwas, Shen Huang |
INTERSPEECH | 2 |
| 2022 | Linguistic-Acoustic Similarity Based Accent Shift for Accent RecognitionabstractGeneral accent recognition (AR) models tend to directly extract low-level information from spectrums, which always significantly overfit on speakers or channels. Considering accent can be regarded as a series of shifts relative to native pronunciation, distinguishing accents will be an easier task with accent shift as input. But due to the lack of native utterance as an anchor, estimating the accent shift is difficult. In this paper, we propose linguistic-acoustic similarity based accent shift (LASAS) for AR tasks. For an accent speech utterance, after mapping the corresponding text vector to multiple accent-associated spaces as anchors, its accent shift could be estimated by the similarities between the acoustic embedding and those anchors. Then, we concatenate the accent shift with a dimension-reduced text vector to obtain a linguistic-acoustic bimodal representation. Compared with pure acoustic embedding, the bimodal representation is richer and more clear by taking full advantage of both linguistic and acoustic information, which can effectively improve AR performance. Experiments on Accented English Speech Recognition Challenge (AESRC) dataset show that our method achieves 77.42% accuracy on Test set, obtaining a 6.94% relative improvement over a competitive system in the challenge. Qijie Shao, Jinghao Yan, Jian Kang 0006, Xian Shi, Pengfei Hu 0004, Lei Xie 0001 |
INTERSPEECH | 6 |
| 2022 | MFA-Conformer: Multi-scale Feature Aggregation Conformer for Automatic Speaker VerificationabstractIn this paper, we present Multi-scale Feature Aggregation Conformer (MFA-Conformer), an easy-to-implement, simple but effective backbone for automatic speaker verification based on the Convolution-augmented Transformer (Conformer).The architecture of the MFA-Conformer is inspired by recent stateof-the-art models in speech recognition and speaker verification.Firstly, we introduce a convolution subsampling layer to decrease the computational cost of the model.Secondly, we adopt Conformer blocks which combine Transformers and convolution neural networks (CNNs) to capture global and local features effectively.Finally, the output feature maps from all Conformer blocks are concatenated to aggregate multi-scale representations before final pooling.We evaluate the MFA-Conformer on the widely used benchmarks.The best system obtains 0.64%, 1.29% and 1.63% EER on VoxCeleb1-O, SITW.Dev, and SITW.Eval set, respectively.MFA-Conformer significantly outperforms the popular ECAPA-TDNN systems in both recognition performance and inference speed.Last but not the least, the ablation studies clearly demonstrate that the combination of global and local feature learning can lead to robust and accurate speaker embedding extraction.We have also released the code 1 for future comparison. Yang Zhang 0025, Zhiqiang Lv, Pengfei Hu 0004, Zhiyong Wu 0001, Hung-yi Lee, Helen M. Meng |
INTERSPEECH | 5 |
| 2021 | Leveraging Phone Mask Training for Phonetic-Reduction-Robust E2E Uyghur Speech RecognitionabstractIn Uyghur speech, consonant and vowel reduction are often encountered, especially in spontaneous speech with high speech rate, which will cause a degradation of speech recognition performance. To solve this problem, we propose an effective phone mask training method for Conformer-based Uyghur end-to-end (E2E) speech recognition. The idea is to randomly mask off a certain percentage features of phones during model training, which simulates the above verbal phenomena and facilitates E2E model to learn more contextual information. According to experiments, the above issues can be greatly alleviated. In addition, deep investigations are carried out into different units in masking, which shows the effectiveness of our proposed masking unit. We also further study the masking method and optimize filling strategy of phone mask. Finally, compared with Conformer-based E2E baseline without mask training, our model demonstrates about 5.51% relative Word Error Rate (WER) reduction on reading speech and 12.92% on spontaneous speech, respectively. The above approach has also been verified on test-set of open-source data THUYG-20, which shows 20% relative improvements. Pengfei Hu 0004, Jian Kang 0006, Shen Huang |
Interspeech | 2 |
| 2021 | The TNT Team System Descriptions of Cantonese and Mongolian for IARPA OpenASR20
Zhiqiang Lv, Ambyer Han, Guan-Bo Wang, Gui-Xin Shi, Jian Kang 0006, Jinghao Yan, Pengfei Hu 0004, Shen Huang, Weiqiang Zhang 0001 |
Interspeech | 8 |
| 2019 | Multimedia Simultaneous Translation System for Minority Language Communication with Mandarin
Shen Huang, Bojie Hu, Pengfei Hu 0004, Jian Kang 0006, Zhiqiang Lv, Jinghao Yan, Qi Ju 0002, Shiyin Kang, Deyi Tuo, Guangzhi Li, Nurmemet Yolwas |
INTERSPEECH | 4 |