VLDB 2026 Research / reviewers in the wild / expert
Yabin Wang 0001
dblp:38/2418-1
· DBLP profile ↗
16ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0003-2931-572XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sample-Aware Knowledge Association and Enhancement for Open-Vocabulary Continual Learning
Zhilin Zhu 0001, Zhiheng Ma, Yabin Wang 0001, Yaguang Song, Yaowei Wang 0001, Xiaopeng Hong |
Int. J. Comput. Vis. | 3 |
| 2026 | Penny-Wise and Pound-Foolish in AI-Generated Image DetectionabstractThe rise of AI-generated images has sparked serious concerns about their potential misuse across various domains, prompting the urgent need for robust detection methods. Despite advancements, many current approaches prioritize short-term gains at the expense of long-term effectiveness. This paper critiques the overly specialized approach of fine-tuning pre-trained models for short-term gains on a single AI image dataset, while disregarding the long-term imperative of achieving generalization and knowledge retention. To address this trade-off issue, we propose a novel learning framework (PoundNet) for the generalization of AI-generated image detection on a pre-trained vision-language model. PoundNet incorporates a learnable prompt design and a balanced objective to preserve broad knowledge from upstream tasks (object classification) while enhancing generalization for downstream tasks (AI-generated image detection). We train PoundNet on a single standard AI image dataset, following common practice in the literature. We then evaluate its performance across 10 large-scale public AI-generated image detection datasets with 5 main evaluation metrics, forming the largest benchmark test set for assessing the generalization ability of AI-generated image detection models, to our knowledge. The comprehensive benchmark evaluation demonstrates that PoundNet successfully balances generalization with knowledge retention, achieving a remarkable relative improvement of 19% in AI-generated image detection performance compared to state-of-the-art methods, while maintaining a strong performance of 63% on object classification tasks. Yabin Wang 0001, Zhiwu Huang, Zhou Su 0001, Adam Prügel-Bennett, Xiaopeng Hong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Dual-Attention based prompt generation and catalyzing for instance-wise continual learning
Xiaopeng Hong, Yabin Wang 0001, Zhiheng Ma, Jinfeng Yang, Dongmei Jiang, Yaowei Wang 0001 |
Pattern Recognit. | 3 |
| 2026 | Linguistic profiling of deepfakes: An open database for next-Generation deepfake detection
Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhiwu Huang |
Pattern Recognit. | 1 |
| 2026 | Asymmetric modal fusion for multi-modal crowd counting
Xiaopeng Hong, Zhiheng Ma, Yabin Wang 0001 |
Pattern Recognit. | 4 |
| 2026 | Continual Conceptual Entity Learning for Text-to-Image Generative ModelsabstractCurrent Text-to-Image generative models struggle to continuously learn multiple distinct entities or concepts, limiting their scalability and hindering practical deployment in dynamic environments. We formulate this task as Continual Conceptual Entity Learning (CEL) and propose a novel framework called Continual Entity Adapter Learning (CEAL). CEAL leverages a compact set of tunable parameters, termed SuperLoRA, to efficient and scalable learning of new entities. We propose a dynamic rank-increasing strategy to train the SuperLoRA, balancing computational efficiency with performance. To evaluate our method, we create three benchmarks encompassing generic objects, human faces, and artistic styles. Experimental results demonstrate that CEAL effectively learns new entities while preserving prior knowledge, outperforming existing methods in both entity fidelity and parameter efficiency. Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhou Su 0001, Zhiwu Huang |
IEEE Trans. Multim. | 1 |
| 2025 | OpenSDI: Spotting Diffusion-Generated Images in the Open WorldabstractThis paper identifies OpenSDI, a challenge for spotting diffusion-generated images in open-world settings. In response to this challenge, we define a new benchmark, the OpenSDI dataset (OpenSDID), which stands out from existing datasets due to its diverse use of large vision-language models that simulate open-world diffusion-based manipulations. Another outstanding feature of OpenSDID is its inclusion of both detection and localization tasks for images manipulated globally and locally by diffusion models. To address the OpenSDI challenge, we propose a Synergizing Pretrained Models (SPM) scheme to build up a mixture of foundation models. This approach exploits a collaboration mechanism with multiple pretrained foundation models to enhance generalization in the OpenSDI context, moving beyond traditional training by synergizing multiple pretrained models through prompting and attending strategies. Building on this scheme, we introduce MaskCLIP, an SPM-based model that aligns Contrastive Language-Image Pre-Training (CLIP) with Masked Autoencoder (MAE). Extensive evaluations on OpenSDID show that MaskCLIP significantly outperforms current state-of-the-art methods for the OpenSDI challenge, achieving remarkable relative improvements of 14.23% in IoU (14.11% in F1) and 2.05% in accuracy (2.38% in F1) compared to the second-best model in localization and detection tasks, respectively. Our dataset and code are available at https://github.com/iamwangyabin/OpenSDI. Yabin Wang 0001, Zhiwu Huang, Xiaopeng Hong |
CVPR | 1 |
| 2025 | Energy Efficient Multi-Robot Task Allocation Constrained by Time Window and PrecedenceabstractTo meet the demands in terms of energy-efficient and fast production and delivery of goods, robotic fleets began to populate warehouses and industrial environments. To maximize the profitability of the operations, multi-robot systems are required to coordinate agents and avoid downtime efficiently. In this paper, agent coordination is formulated as a multi-robot task allocation (MRTA) problem with time and precedence constraints. The method capitalizes on a graph method to build a measure graph reflecting the sparsity of tasks and a precedence graph, which includes the task constraints, to group the tasks into batches. A batch solver is provided to obtain the final solutions to the MRTA. In this way, the sustainability and environmental impact of logistics operations can be improved by reducing the number of robots needed to complete tasks and also by assigning tasks closest to the robot location, reducing the amount of time and the total energy required for the robots to complete the job. Extensive experiments on both uniformly distributed and sparse data sets prove the effectiveness of the proposed algorithm compared to state-of-the-art algorithms such as MIP and TePSSI.Note to Practitioners—This paper was motivated by the problem of minimizing the energy consumption of multi-robot systems in the execution of complex tasks, which requires, in the most general case, the motion of the robot to a target location and further on-site operations. This scenario is particularly relevant in smart, automated warehouses, where mobile robots are repeatedly demanded to store or dispatch goods in a structured environment, where operation duration and future requests are known a priori. The paper formulates this problem by means of a batched multi-robot task allocation (BMRTA) optimization, which can include time windows and precedence constraints jointly. First, the task constraints are encoded into two graphs and then combined to group subtasks together in batches. Then, each batch is solved separately, minimizing the overall energy required to achieve the tasks in the batch. Although the optimality of the solution is ensured only locally, i.e., within the same batch, the task clustering improves the computational efficiency with respect to global approaches, especially in large-sized problems. Experimental results demonstrated that when comparing BMRTA with literature approaches such as MIP and TePSSI, not only the energy consumption but also the total travel distance can be minimized while the total duration of the tasks remains comparable. Lixuan Zhang, Jianzhuang Zhao, Edoardo Lamon, Yabin Wang 0001, Xiaopeng Hong |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Multi-modal Crowd Counting via Modal Emulation
Xiaopeng Hong, Zhiheng Ma, Yabin Wang 0001, Xiaopeng Fan 0001 |
BMVC | 5 |
| 2024 | Token-based deep reinforcement learning for Heterogeneous VRP with Service Time Constraints
Xiaopeng Hong, Yabin Wang 0001, Junzhou Zhao, Guanghui Sun, Baoxing Qin |
Knowl. Based Syst. | 3 |
| 2023 | Isolation and Impartial Aggregation: A Paradigm of Incremental Learning without InterferenceabstractThis paper focuses on the prevalent stage interference and stage performance imbalance of incremental learning. To avoid obvious stage learning bottlenecks, we propose a new incremental learning framework, which leverages a series of stage-isolated classifiers to perform the learning task at each stage, without interference from others. To be concrete, to aggregate multiple stage classifiers as a uniform one impartially, we first introduce a temperature-controlled energy metric for indicating the confidence score levels of the stage classifiers. We then propose an anchor-based energy self-normalization strategy to ensure the stage classifiers work at the same energy level. Finally, we design a voting-based inference augmentation strategy for robust inference. The proposed method is rehearsal-free and can work for almost all incremental learning scenarios. We evaluate the proposed method on four large datasets. Extensive results demonstrate the superiority of the proposed method in setting up new state-of-the-art overall performance. Code is available at https://github.com/iamwangyabin/ESN. Yabin Wang 0001, Zhiheng Ma, Zhiwu Huang, Yaowei Wang 0001, Zhou Su 0001, Xiaopeng Hong |
AAAI | 1 |
| 2023 | A Continual Deepfake Detection Benchmark: Dataset, Methods, and EssentialsabstractThere have been emerging a number of benchmarks and techniques for the detection of deepfakes. However, very few works study the detection of incrementally appearing deepfakes in the real-world scenarios. To simulate the wild scenes, this paper suggests a continual deepfake detection benchmark (CDDB) over a new collection of deepfakes from both known and unknown generative models. The suggested CDDB designs multiple evaluations on the detection over easy, hard, and long sequence of deepfake tasks, with a set of appropriate measures. In addition, we exploit multiple approaches to adapt multiclass incremental learning methods, commonly used in the continual visual recognition, to the continual deepfake detection problem. We evaluate existing methods, including their adapted ones, on the proposed CDDB. Within the proposed benchmark, we explore some commonly known essentials of standard continual learning. Our study provides new insights on these essentials in the context of continual deepfake detection. The suggested CDDB is clearly more challenging than the existing benchmarks, which thus offers a suitable evaluation avenue to the future research. Both data and code are available at https://github.com/Coral79/CDDB. Chuqiao Li, Zhiwu Huang, Danda Pani Paudel, Yabin Wang 0001, Mohamad Shahbazi, Xiaopeng Hong, Luc Van Gool |
WACV | 4 |
| 2022 | S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningabstractState-of-the-art deep neural networks are still struggling to address the catastrophic forgetting problem in continual learning. In this paper, we propose one simple paradigm (named as S-Prompting) and two concrete approaches to highly reduce the forgetting degree in one of the most typical continual learning scenarios, i.e., domain increment learning (DIL). The key idea of the paradigm is to learn prompts independently across domains with pre-trained transformers, avoiding the use of exemplars that commonly appear in conventional methods. This results in a win-win game where the prompting can achieve the best for each domain. The independent prompting across domains only requests one single cross-entropy loss for training and one simple K-NN operation as a domain identifier for inference. The learning paradigm derives an image prompt learning approach and a novel language-image prompt learning approach. Owning an excellent scalability (0.03% parameter increase per domain), the best of our approaches achieves a remarkable relative improvement (an average of about 30%) over the best of the state-of-the-art exemplar-free methods for three standard DIL tasks, and even surpasses the best of them relatively by about 6% in average when they use exemplars. Source code is available at https://github.com/iamwangyabin/S-Prompts. Yabin Wang 0001, Zhiwu Huang, Xiaopeng Hong |
NeurIPS | 1 |
| 2022 | ECCNAS: Efficient Crowd Counting Neural Architecture SearchabstractRecent solutions to crowd counting problems have already achieved promising performance across various benchmarks. However, applying these approaches to real-world applications is still challenging, because they are computation intensive and lack the flexibility to meet various resource budgets. In this article, we propose an efficient crowd counting neural architecture search (ECCNAS) framework to search efficient crowd counting network structures, which can fill this research gap. A novel search from pre-trained strategy enables our cross-task NAS to explore the significantly large and flexible search space with less search time and get more proper network structures. Moreover, our well-designed search space can intrinsically provide candidate neural network structures with high performance and efficiency. In order to search network structures according to hardwares with different computational performance, we develop a novel latency cost estimation algorithm in our ECCNAS. Experiments show our searched models get an excellent trade-off between computational complexity and accuracy and have the potential to deploy in practical scenarios with various resource budgets. We reduce the computational cost, in terms of multiply-and-accumulate (MACs), by up to 96% with comparable accuracy. And we further designed experiments to validate the efficiency and the stability improvement of our proposed search from pre-trained strategy. Yabin Wang 0001, Zhiheng Ma, Xing Wei 0001, Yaowei Wang 0001, Xiaopeng Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | A Hardware-adaptive Deep Feature Matching Pipeline for Real-time 3D Reconstruction
Yabin Wang 0001, Baotong Li, Xin Li 0003 |
Comput. Aided Des. | 2 |
| 2020 | Deep memory network with Bi-LSTM for personalized context-aware citation recommendation
Jie Wang 0072, Li Zhu 0003, Tao Dai 0002, Yabin Wang 0001 |
Neurocomputing | 4 |