VLDB 2026 Research / reviewers in the wild / expert
Boqing Zhu
dblp:217/2975
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-7867-2112ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Underwater acoustic signal denoising with diffusion-based generative models
Boqing Zhu, Yanxin Ma, Zemin Zhou, Xiaoqian Zhu |
Signal Process. | 1 |
| 2025 | SSAST-Adapter: A Parameter-efficient Incremental Learning Algorithm for Underwater Acoustic Target RecognitionabstractUnderwater acoustic target recognition involves identifying and classifying targets in underwater environments using acoustic signals. In recent years, deep learning has made significant progress in this field. However, the models require the entire dataset to be available upfront, and classification categories must be predefined. In practical scenarios, new objects or species may appear in underwater environments over time, making it impractical to retrain a model from scratch each time new data is introduced. At the same time, training on new data inevitably leads to catastrophic forgetting of past data. To address these challenges, we propose a large-scale pre-training strategy combined with adapters to enable incremental learning without the need for complete retraining. To minimize the number of parameters during model fine-tuning, we employ a adapter structure at various network layers, reducing the number of trainable parameters to less than 2%. Experimental results demonstrate that our proposed method effectively reduces the need for full retraining by allowing the model to update in a resource-efficient manner using only new data, and the model has achieved high recognition accuracy in underwater acoustic target recognition tasks. Qisheng Xu, Boqing Zhu, Zijian Gao, Lingbin Zeng, Kele Xu |
ICASSP | 3 |
| 2024 | Learning incremental audio-visual representation for continual multimodal understandingabstractDeep learning methods have demonstrated remarkable success in processing static datasets for various video tasks. However, when confronted with continuous data streams, these approaches often encounter the challenge of catastrophic forgetting. This phenomenon leads to a significant decline in overall performance when learning new classes incrementally. Moreover, existing methods tend to overlook the correlation between audio and visual modalities in video incremental learning, despite their joint significance in scene comprehension. How to continuously learn from new classes while maintaining the knowledge of old videos with limited storage and computing resources is becoming imperative in the field of multimodal learning. In this paper, we introduce CavRL, a pioneering benchmark for audio–visual representation learning under class incremental scenarios. To mitigate catastrophic forgetting, we propose a rehearsal-based training approach that leverages a small exemplar set from previous classes. Our approach constrains the memory buffer within strict storage limits, optimizing exemplar selection by learning correlative audio–visual representations. Additionally, we employ a distillation method to mitigate forgetting in a self-supervised manner. Evaluations on two prevalent multimodal tasks: audio–visual event classification and audio–visual speaker recognition, which demonstrate that CavRL outperforms existing state-of-the-art incremental learning methods across various settings. We anticipate that CavRL will significantly advance research in continual multimodal learning. • Catastrophic Forgetting Challenge in Multimodal Learning . Deep learning grapples with catastrophic forgetting in continuous mulitimodal data. • Consideration Audio–Visual Correlation . Consideration relationship between audio and visual modalities in continual learning. • Self-Supervised Distillation to Alleviate Forgetting . CavRL implements a self-supervised distillation method to reduce forgetting. • CavRL Benchmark . CavRL is a new benchmark for audio–visual learning in class incremental scenarios. Boqing Zhu, Kele Xu, Zemin Zhou, Xiaoqian Zhu |
Knowl. Based Syst. | 1 |
| 2024 | SpeedAdv: Enabling Green Light Optimized Speed Advisory for Diverse Traffic LightsabstractGreen Light Optimized Speed Advisory (GLOSA) systems have emerged to allow drivers to pass traffic lights during a green interval. However, various adaptive and intelligent traffic light control approaches have been adopted in many cities, resulting in the development of current GLOSA technologies lagging behind that of traffic light technologies. When taking diverse dynamic traffic lights into account, it is difficult to model the interactions between vehicles and traffic lights, which is further exacerbated by the hybrid control strategies of traffic lights. To this end, we design a new GLOSA systemSpeedAdvto provide optimal speed advisory for addressing diverse traffic lights. We formulate the problem as a Multi-Agent Markov Decision Process (MAMDP) with an implicit common goal and propose a heterogeneous-agent collaborative framework based on reinforcement learning. Three main modules are used in the system: i) a spatio-temporal relation reasoning module based on the phase-aware attention mechanism pays more attention to the traffic rules and traffic flow diversion of adjacent intersections to predict traffic conditions for a few seconds later; ii) a behavior approximating module based on imitation learning is introduced to approximate the phases of diverse traffic lights; iii) a speed advisory module provides the optimal speed advisory based on policy gradient reinforcement learning relying on the above two modules and other information collected by vehicles. We implement and evaluateSpeedAdvwith a real-world trajectory dataset, together with a field test based on a prototype system, demonstrating thatSpeedAdvimproves the overall performance by at least 24.1% in terms of travel time, energy consumption, safety, and comfort compared to the state-of-the-artGreenDrivemethod. Lige Ding, Dong Zhao 0001, Boqing Zhu, Zhaofeng Wang, Jianjun Tong, Huadong Ma |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Raw Ultrasound-Based Phonetic Segments Classification Via Mask ModelingabstractUltrasound tongue imaging is widely used in clinical linguistics and phonetics. Recently, deep neural networks, especially convolutional neural networks, have been widely used in the interpretation and analysis of ultrasound tongue images (UTI). Despite achieving satisfactory performance, deep models rely on a large amount of manually labeled data, which is often difficult to obtain in practical settings. To address this issue, this paper focuses on how to utilize a large amount of unlabeled UTI data to improve the performance of UTI classification task. Specifically, we explore self-supervised learning with masking modeling strategy. By predicting the masked part, our pre-trained model enables the neural network to infer contextual information. Then, we fine-tune the pre-trained model with a small amount of labeled data. Compared with the previous competing algorithms, our method can improve the classification accuracy by an average of 13.33% in four different scenarios. Kang You, Bo Liu 0014, Kele Xu, Yunsheng Xiong, Qisheng Xu, Ming Feng, Tamás Gábor Csapó, Boqing Zhu |
ICASSP | 8 |
| 2023 | Multi-task Pre-training Language Model for Semantic Network CompletionabstractSemantic networks, exemplified by the knowledge graph, serve as a means to represent knowledge by leveraging the structure of a graph. While the knowledge graph exhibits promising potential in the field of natural language processing, it suffers from incompleteness. This article focuses on the task of completing knowledge graphs by predicting linkages between entities, which is fundamental yet critical. Traditional methods based on translational distance struggle when dealing with unseen entities. In contrast, semantic matching presents itself as a potential solution due to its ability to handle such cases. However, semantic matching-based approaches necessitate large-scale datasets for effective training, which are typically unavailable in practical scenarios, hindering their competitive performance. To address this challenge, we propose a novel architecture for knowledge graphs known as LP-BERT, which incorporates a language model. LP-BERT consists of two primary stages: multi-task pre-training and knowledge graph fine-tuning. During the pre-training phase, the model acquires relationship information from triples by predicting either entities or relations through three distinct tasks. In the fine-tuning phase, we introduce a batch-based triple-style negative sampling technique inspired by contrastive learning. This method significantly increases the proportion of negative sampling while maintaining a nearly unchanged training time. Furthermore, we propose a novel data augmentation approach that leverages the inverse relationship of triples to enhance both the performance and robustness of the model. To demonstrate the effectiveness of our proposed framework, we conduct extensive experiments on three widely used knowledge graph datasets: WN18RR, FB15k-237, and UMLS. The experimental results showcase the superiority of our methods, with LP-BERT achieving state-of-the-art performance on the WN18RR and FB15k-237 datasets. Boqing Zhu, Sen Yang 0003, Kele Xu, Yukai He, Huaimin Wang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype ContrastabstractWe present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice and face. These methods neglect the semantic content of different videos, introducing false-negative pairs as training noise. Furthermore, the positive pairs are constructed based on the natural correlation between audio clips and visual frames. However, this correlation might be weak or inaccurate in a large amount of real-world data, which leads to deviating positives into the contrastive paradigm. To address these issues, we propose the cross-modal prototype contrastive learning (CMPC), which takes advantage of contrastive methods and resists adverse effects of false negatives and deviate positives. On one hand, CMPC could learn the intra-class invariance by constructing semantic-wise positives via unsupervised clustering in different modalities. On the other hand, by comparing the similarities of cross-modal instances from that of cross-modal prototypes, we dynamically recalibrate the unlearnable instances' contribution to overall loss. Experiments show that the proposed approach outperforms state-of-the-art unsupervised methods on various voice-face association evaluation protocols. Additionally, in the low-shot supervision setting, our method also has a significant improvement compared to previous instance-wise contrastive learning. Boqing Zhu, Kele Xu, Zheng Qin 0002, Tao Sun 0005, Huaimin Wang 0001, Yuxing Peng 0001 |
IJCAI | 1 |
| 2022 | Multiple Temporal Fusion based Weakly-supervised Pre-training Techniques for Video CategorizationabstractIn this paper, we present our solution of the ACM Multimedia 2022 pre-training for video understanding challenge. First, we pre-train the models on large-scale weakly-supervised video datasets with different temporal resolutions, then fine-tune the model for downstream application. Quantitative comparisons are conducted to evaluate the performance of different networks at multiple temporal resolutions. Moreover, we fusion different pre-trained models through weighted averaging. We achieve an accuracy of 62.39% in the testing set, which ranked as the first place in the video categorization track of this challenge. Xiaochen Cai, Hengxing Cai, Boqing Zhu, Kele Xu, Wei-Wei Tu |
ACM Multimedia | 3 |
| 2022 | Masked Modeling-based Audio Representation for ACM Multimedia 2022 Computational Paralinguistics ChallengEabstractIn this paper, we present our solution for ACM Multimedia 2022 Computational Paralinguistics Challenge. Our method employs the self-supervised learning paradigm, as it achieves promising results in computer vision and audio signal processing. Specifically, we firstly explore modifying the Swin Transformer architecture to learn general representation for the audio signals, accompanied with random masking on the log-mel spectrogram. The main goal of the pretext task is to predict the masked parts, by combining the advantages of the Swin-Transformer and masked modeling. For the downstream tasks, we utilize the labelled datasets to fine-tune the pre-trained model. Compared with the competitive baselines, our approach can provide significant performance improvements without ensembling. Kang You, Kele Xu, Boqing Zhu, Ming Feng, Bo Liu 0014, Bo Ding 0001 |
ACM Multimedia | 3 |
| 2020 | Audio Tagging by Cross Filtering Noisy LabelsabstractHigh quality labeled datasets have allowed deep learning to achieve impressive results on many sound analysis tasks. Yet, it is labor-intensive to accurately annotate large amount of audio data, and the dataset may contain noisy labels in the practical settings. Meanwhile, the deep neural networks are susceptive to those incorrect labeled data because of their outstanding memorization ability. In this article, we present a novel framework, named CrossFilter, to combat the noisy labels problem for audio tagging. Multiple representations (such as, Logmel and MFCC) are used as the input of our framework for providing more complementary information of the audio. Then, though the cooperation and interaction of two neural networks, we divide the dataset into curated and noisy subsets by incrementally pick out the possibly correctly labeled data from the noisy data. Moreover, our approach leverages the multi-task learning on curated and noisy subsets with different loss function to fully utilize the entire dataset. The noisy-robust loss function is employed to alleviate the adverse effects of incorrect labels. On both the audio tagging datasets FSDKaggle2018 and FSDKaggle2019, empirical results demonstrate the performance improvement compared with other competing approaches. On FSDKaggle2018 dataset, our method achieves state-of-the-art performance and even surpasses the ensemble models. Boqing Zhu, Kele Xu, Qiuqiang Kong, Huaimin Wang 0001, Yuxing Peng 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | A Mobile Application for Sound Event DetectionabstractSound event detection is intended to analyze and recognize the sound events in audio streams and it has widespread applications in real life. Recently, deep neural networks such as convolutional recurrent neural networks have shown state-of-the-art performance in this task. However, the previous methods were designed and implemented on devices with rich computing resources, and there are few applications on mobile devices. This paper focuses on the solution on the mobile platform for sound event detection. The architecture of the solution includes offline training and online detection. During offline training process, multi model-based distillation method is used to compress model to enable real-time detection. The online detection process includes acquisition of sensor data, processing of audio signals, and detecting and recording of sound events. Finally, we implement an application on the mobile device that can detect sound events in near real time. Yingwei Fu, Kele Xu, Haibo Mi, Huaimin Wang 0001, Boqing Zhu |
IJCAI | 6 |
| 2018 | Multi-LCNN: A Hybrid Neural Network Based on Integrated Time-Frequency Characteristics for Acoustic Scene ClassificationabstractAcoustic scene classification (ASC) is an important task in audio signal processing and can be useful in many real-world applications. Recently, several deep neural network models have been proposed for ASC, such as LSTMs based on temporal analysis and CNNs based on frequency spectrum, as well as hybrid models of LSTM and CNN to further improve classification performance. However, existing hybrid models fail to properly preserve the temporal information when transferring data between different models. In this work, we first analyze the cause of such temporal information loss. We then propose Multi-LCNN, a new hybrid model with two important mechanisms: (1) a LCNN architecture to effectively preserve temporal information; and (2) a multi-channel feature fusion mechanism (MCFF) that combines enhanced temporal information and frequency spectrogram information to learn highly integrated and discriminative features for ASC. Evaluations on the TUT ASC 2016 dataset show that our model can achieve an improvement of 10.23% over the baseline method, and is currently the best-performing end-to-end model on this dataset. Jin Lei, Boqing Zhu, Qin Lv, Zhen Huang 0006, Yuxing Peng 0001 |
ICTAI | 3 |
| 2018 | Learning Environmental Sounds with Multi-scale Convolutional Neural NetworkabstractDeep learning has dramatically improved the performance of sounds recognition. However, learning acoustic models directly from the raw waveform is still challenging. Current waveform-based models generally use time-domain convolutional layers to extract features. The features extracted by single size filters are insufficient for building discriminative representation of audios. In this paper, we propose multi-scale convolution operation, which can get better audio representation by improving the frequency resolution and learning filters cross all frequency area. For leveraging the waveform-based features and spectrogram-based features in a single model, we introduce two-phase method to fuse the different features. Finally, we propose a novel end-to-end network called WaveMsNet based on the multi-scale convolution operation and two-phase method. On the environmental sounds classification datasets ESC-10 and ESC-50, the classification accuracies of our WaveMsNet achieve 93.75% and 79.10% respectively, which improve significantly from the previous methods. Boqing Zhu, Jin Lei, Zhen Huang 0006, Yuxing Peng 0001, Fei Li 0004 |
IJCNN | 1 |