EDBT 2026 Demo / reviewers in the wild / expert
Boquan Li 0002
dblp:116/7306-2
· DBLP profile ↗
14ranked-venue papers
1as first author
13since 2021 · last 2026
0009-0002-0865-7368ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Be Responsible in Your Answers! Monitoring Out-of-Domain Behaviors in Domain-Specific LLMs
Boquan Li 0002, Chenzhe Lou, Zhe Ren, Peixin Zhang 0001, Zirui Fu, Jun Sun 0001, Yaowen Zheng |
WWW | 1 |
| 2026 | SpeakerMatch: Matching reliable pseudo-labels in semi-supervised and self-supervised speaker recognition with confidence distributionabstractFor speaker recognition, pseudo-labeling has shown advantages in alleviating the scarcity of labeled data. Inspired by image classification tasks, existing methods typically adopt threshold-based strategies to identify reliable pseudo-labels. However, compared to image classification, speaker recognition requires finer-grained class discrimination for open-set identity verification, and thus often adopts margin-based losses that amplify gradients near decision boundaries. While effective under full supervision, this design increases sensitivity to noisy or sparse pseudo-labels, limiting the effectiveness of threshold-based selection. In this work, we propose SpeakerMatch , a novel distribution-based framework for semi-supervised and self-supervised speaker recognition. SpeakerMatch models the confidence distribution to distinguish reliable pseudo-labels from noisy ones globally and selects those whose confidence values and confidence prediction behaviors closely align with high-quality signals. Systematic evaluation across five settings shows that our method outperforms existing approaches, achieving a 13.7% relative improvement over the best semi-supervised speaker recognition baseline, while also delivering a lower equal error rate (EER) and reduced training costs compared to self-supervised methods. Jinghan Liu, Xingmei Wang 0002, Jiaxiang Meng, Boquan Li 0002 |
Signal Process. | 4 |
| 2025 | Int*-Match: Balancing Intra-Class Compactness and Inter-Class Discrepancy for Semi-Supervised Speaker RecognitionabstractOpen-set speaker recognition is to identify whether the voices are from the same speaker. One challenge of speaker recognition is collecting large amounts of high-quality data. Based on the promising results of image classification, one intuitively feasible solution is semi-supervised learning (SSL) which uses confidence thresholds to assign pseudo labels for unlabeled data. However, we empirically demonstrated that applying SSL methods to speaker recognition is non-trivial. These methods focus solely on inter-class discrepancy as thresholds to select pseudo labels, overlooking intra-class compactness, which is particularly important for open-set speaker recognition tasks. Motivated by this, we propose Int*-Match, a semi-supervised speaker recognition method selecting reliable pseudo labels with intra-class compactness and inter-class discrepancy for speaker recognition. In particular, we use the inter-class discrepancy of labeled data as the threshold for pseudo-label selection and adjust the threshold based on the intra-class compactness of the pseudo labels dynamically and adaptively. Our systematic experiments demonstrate the superiority of Int*-Match, presenting an outstanding Equal Error Rate (EER) of 1.00% on the VoxCeleb1 original test set, which is merely 0.06% below the performance achieved by fully supervised learning. Xingmei Wang 0002, Jinghan Liu, Jiaxiang Meng, Boquan Li 0002 |
AAAI | 4 |
| 2024 | Assessing Backdoor Risk in Deepfake Detection
Boquan Li 0002, Min Yu 0001, Kam-Pui Chow, Fuqiang Du, Weiqing Huang |
IFIP Int. Conf. Digital Forensics | 2 |
| 2024 | Two-stage Semi-supervised Speaker Recognition with Gated Label Learning
Xingmei Wang 0002, Jiaxiang Meng, Kong-Aik Lee, Boquan Li 0002, Jinghan Liu |
IJCAI | 4 |
| 2024 | SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery DetectionabstractDetection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leveraging the synergy between audio and visual speech elements, embarking on a novel approach through audio-visual speech representation learning. Our work is motivated by the finding that audio signals, enriched with speech content, can provide precise information effectively reflecting facial movements. To this end, we first learn precise audio-visual speech representations on real videos via a self-supervised masked prediction task, which encodes both local and global semantic information simultaneously. Then, the derived model is directly transferred to the forgery detection task. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods in terms of cross-dataset generalization and robustness, without the participation of any fake video in model training. Yachao Liang, Min Yu 0001, Gang Li 0009, Boquan Li 0002, Weiqing Huang |
NeurIPS | 5 |
| 2024 | CATNet: Cross-modal fusion for audio-visual speech recognition
Xingmei Wang 0002, Jiachen Mi, Boquan Li 0002, Yixu Zhao, Jiaxiang Meng |
Pattern Recognit. Lett. | 3 |
| 2023 | Hierarchical memory-guided long-term tracking with meta transformer inquiry network
Xingmei Wang 0002, Guohao Nie, Boquan Li 0002, Minyang Kang |
Knowl. Based Syst. | 3 |
| 2021 | Adaptive Smooth L1 Loss: A Better Way to Regress Scene Texts with Extreme Aspect RatiosabstractIn recent years, scene text detection has experienced rapid development. Regression-based methods are currently a mainstream method for scene text detection, and the effect of bounding box regression is a major factor limiting their detection performance. The regression of bounding boxes is greatly affected by the aspect ratio of texts since the text in natural scenes varies greatly in height and width. However, the existing methods ignore the difference between the height and width of the text in the bounding box regression, which leads to an imperfect regression effect and thus suppresses the performance of the scene text detection. In this paper, we propose an Adaptive Smooth L1 Loss function (abbreviated as ASLL) for bounding box regression, which can adaptively determine the weight of each regression variable according to the current state of the model during the training process, so as to guide the bounding box to regress in a more critical direction. The experimental results demonstrate that ASLL achieves promising performance on scene text detection. Specially, an F-measure of 84.56% is achieved on CTW-1500 dataset, surpassing the state-of-the-art detectors, and the detection results on TotalText and ICDAR2015 datasets are competitive to those of state-of-the-art methods. Chao Liu 0020, Min Yu 0001, Baole Wei, Boquan Li 0002, Gang Li 0009, Weiqing Huang |
ISCC | 5 |
| 2021 | Finding disposable domain names: A linguistics-based stacking approach
Yuwei Zeng, Xiao-chun Yun, Xunxun Chen, Boquan Li 0002, Haiwei Tsang, Yipeng Wang 0001, Tianning Zang, Yongzheng Zhang 0002 |
Comput. Networks | 4 |
| 2021 | An end-to-end text spotter with text relation networksabstractAbstract Reading text in images automatically has become an attractive research topic in computer vision. Specifically, end-to-end spotting of scene text has attracted significant research attention, and relatively ideal accuracy has been achieved on several datasets. However, most of the existing works overlooked the semantic connection between the scene text instances, and had limitations in situations such as occlusion, blurring, and unseen characters, which result in some semantic information lost in the text regions. The relevance between texts generally lies in the scene images. From the perspective of cognitive psychology, humans often combine the nearby easy-to-recognize texts to infer the unidentifiable text. In this paper, we propose a novel graph-based method for intermediate semantic features enhancement, called Text Relation Networks. Specifically, we model the co-occurrence relationship of scene texts as a graph. The nodes in the graph represent the text instances in a scene image, and the corresponding semantic features are defined as representations of the nodes. The relative positions between text instances are measured as the weights of edges in the established graph. Then, a convolution operation is performed on the graph to aggregate semantic information and enhance the intermediate features corresponding to text instances. We evaluate the proposed method through comprehensive experiments on several mainstream benchmarks, and get highly competitive results. For example, on the , our method surpasses the previous top works by 2.1% on the word spotting task. Baole Wei, Min Yu 0001, Gang Li 0009, Boquan Li 0002, Chao Liu 0020, Weiqing Huang |
Cybersecur. | 5 |
| 2021 | FakeFilter: A cross-distribution Deepfake detection system with domain adaptationabstractAbuse of face swap techniques poses serious threats to the integrity and authenticity of digital visual media. More alarmingly, fake images or videos created by deep learning technologies, also known as Deepfakes, are more realistic, high-quality, and reveal few tampering traces, which attracts great attention in digital multimedia forensics research. To address those threats imposed by Deepfakes, previous work attempted to classify real and fake faces by discriminative visual features, which is subjected to various objective conditions such as the angle or posture of a face. Differently, some research devises deep neural networks to discriminate Deepfakes at the microscopic-level semantics of images, which achieves promising results. Nevertheless, such methods show limited success as encountering unseen Deepfakes created with different methods from the training sets. Therefore, we propose a novel Deepfake detection system, named FakeFilter, in which we formulate the challenge of unseen Deepfake detection into a problem of cross-distribution data classification, and address the issue with a strategy of domain adaptation. By mapping different distributions of Deepfakes into similar features in a certain space, the detection system achieves comparable performance on both seen and unseen Deepfakes. Further evaluation and comparison results indicate that the challenge has been successfully addressed by FakeFilter. Boquan Li 0002, Baole Wei, Gang Li 0009, Chao Liu 0020, Weiqing Huang, Meimei Li, Min Yu 0001 |
J. Comput. Secur. | 2 |
| 2021 | Cloud-Based Data Offloading for Multi-focus and Multi-views Image Fusion in Mobile Applications
Yiqi Shi, Liang Kou, Boquan Li 0002, Qing Yang 0003, Liguo Zhang 0002 |
Mob. Networks Appl. | 5 |
| 2019 | Restoration as a Defense Against Adversarial Perturbations for Spam Image Detection
Boquan Li 0002, Min Yu 0001, Chao Liu 0020, Weiqing Huang, Lejun Fan, Jianfeng Xia |
ICANN (3) | 2 |