EDBT 2026 Demo / reviewers in the wild / expert
Wenbin An
dblp:331/2394
· DBLP profile ↗
16ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0003-0062-7201ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Retrieval-Augmented Large Vision Language Models via Knowledge Conflict MitigationabstractMultimodal Retrieval-Augmented Generation (MRAG) has recently been explored to empower Large Vision Language Models (LVLMs) with more comprehensive and up-to-date contextual knowledge, aiming to compensate for their limited and coarse-grained parametric knowledge in knowledge-intensive tasks. However, the retrieved contextual knowledge is usually not aligned with LVLMs’ internal parametric knowledge, leading to knowledge conflicts and further unreliable responses. To tackle this issue, we design KCM, a training-free and plug-and-play framework that can effectively mitigate knowledge conflicts while incorporating MRAG for more accurate LVLM responses. KCM enhances contextual knowledge utilization by modifying the LVLM architecture from three key perspectives. First, KCM adaptively adjusts attention distributions among multiple attention heads, encouraging LVLMs to focus on contextual knowledge with reduced distraction. Second, KCM identifies and prunes knowledge-centric LVLM neurons that encode coarse-grained parametric knowledge, thereby suppressing interferences and enabling more effective integration of contextual knowledge. Third, KCM amplifies the information flow from the input context by injecting supplementary context logits, reinforcing its contribution to the final output. Extensive experiments over multiple LVLMs and benchmarks show that KCM outperforms the state-of-the-art consistently by large margins, incurring neither extra training nor external tools. Wenbin An, Jiahao Nie 0002, Feng Tian 0002, Mingxiang Cai, Yaqiang Wu, Shijian Lu |
AAAI | 1 |
| 2026 | Graph Mixture of Experts with Differential Cross-Attention Alignment for Multimodal Intent Recognition
Shilin Sun 0001, Wenbin An, Qidong Liu 0002, Jiahao Nie 0002, Zhi Zeng 0001, Xian-Sheng Hua, Yaqiang Wu, Feng Tian 0002 |
Knowl. Based Syst. | 2 |
| 2025 | Unleashing the Potential of Model Bias for Generalized Category DiscoveryabstractGeneralized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another labeled dataset containing only known categories. The primary challenges stem from model bias induced by pre-training on only known categories and the lack of precise supervision for novel ones, leading to category bias towards known categories and category confusion among different novel categories, which hinders models' ability to identify novel categories effectively. To address these challenges, we propose a novel framework named Self-Debiasing Calibration (SDC). Unlike prior methods that regard model bias towards known categories as an obstacle to novel category identification, SDC provides a novel insight into unleashing the potential of the bias to facilitate novel category learning. Specifically, we utilize the biased pre-trained model to guide the subsequent learning process on unlabeled data. The output of the biased model serves two key purposes. First, it provides an accurate modeling of category bias, which can be utilized to measure the degree of bias and debias the output of the current training model. Second, it offers valuable insights for distinguishing different novel categories by transferring knowledge between similar categories. Based on these insights, SDC dynamically adjusts the output logits of the current training model using the output of the biased model. This approach produces less biased logits to effectively address the issue of category bias towards known categories, and generates more accurate pseudo labels for unlabeled data, thereby mitigating category confusion for novel categories. Experiments on three benchmark datasets show that SDC outperforms SOTA methods, especially in the identification of novel categories. Wenbin An, Haonan Lin, Jiahao Nie 0002, Feng Tian 0002, Wenkai Shi, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 1 |
| 2025 | Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionabstractDespite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsistent with the actual objects in images. We examine different LVLMs and pinpoint that one root cause of object hallucinations lies with deficient attention on discriminative image features. Specifically, LVLMs often predominantly attend to prompt-irrelevant global features instead of prompt-relevant local features, undermining their visual grounding capacity and leading to object hallucinations. We propose Assembly of Global and Local Attention (AGLA), a training-free and plug-and-play approach that mitigates hallucinations by assembling global features for response generation and local features for visual discrimination simultaneously. Specifically, we introduce an image-prompt matching scheme that captures prompt-relevant local features from images, leading to an augmented view of the input image where prompt-relevant content is highlighted while irrelevant distractions are suppressed. Hallucinations can thus be mitigated with a calibrated logit distribution that is from generative global features of the original image and discriminative local features of the augmented image. Extensive experiments show the superiority of AGLA in LVLM hallucination mitigation, demonstrating its wide applicability across both discriminative and generative tasks. Our code is available at https://github.com/Lackel/AGLA. Wenbin An, Feng Tian 0002, Sicong Leng, Jiahao Nie 0002, Haonan Lin, Qianying Wang 0002, Ping Chen 0001, Shijian Lu |
CVPR | 1 |
| 2025 | Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention ReallocationabstractDespite their recent progress, Multimodal Large Language Models (MLLMs) often struggle in knowledge-intensive tasks due to the limited and outdated parametric knowledge acquired during training. Multimodal Retrieval Augmented Generation addresses this issue by retrieving contextual knowledge from external databases, thereby enhancing MLLMs with expanded knowledge sources.
However, existing MLLMs often fail to fully leverage the retrieved contextual knowledge for response generation. We examine representative MLLMs and identify two major causes, namely, attention bias toward different tokens and knowledge conflicts between parametric and contextual knowledge. To this end, we design Adaptive Logits Fusion and Attention Reallocation (ALFAR), a training-free and plug-and-play approach that improves MLLM responses by maximizing the utility of the retrieved knowledge. Specifically, ALFAR tackles the challenges from two perspectives. First, it alleviates attention bias by adaptively shifting attention from visual tokens to relevant context tokens according to query-context relevance. Second, it decouples and weights parametric and contextual knowledge at output logits, mitigating conflicts between the two types of knowledge. As a plug-and-play method, ALFAR achieves superior performance across diverse datasets without requiring additional training or external tools. Extensive experiments over multiple MLLMs and benchmarks show that ALFAR consistently outperforms the state-of-the-art by large margins. Our code and data are available at https://github.com/Lackel/ALFAR. Wenbin An, Jiahao Nie 0002, Feng Tian 0002, Haonan Lin, Mingxiang Cai, Yaqiang Wu, Qianying Wang 0002, Shijian Lu |
NeurIPS | 1 |
| 2024 | Transfer and Alignment Network for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is a crucial real-world task that aims to recognize both known and novel categories from an unlabeled dataset by leveraging another labeled dataset with only known categories. Despite the improved performance on known categories, current methods perform poorly on novel categories. We attribute the poor performance to two reasons: biased knowledge transfer between labeled and unlabeled data and noisy representation learning on the unlabeled data. The former leads to unreliable estimation of learning targets for novel categories and the latter hinders models from learning discriminative features. To mitigate these two issues, we propose a Transfer and Alignment Network (TAN), which incorporates two knowledge transfer mechanisms to calibrate the biased knowledge and two feature alignment mechanisms to learn discriminative features. Specifically, we model different categories with prototypes and transfer the prototypes in labeled data to correct model bias towards known categories. On the one hand, we pull instances with known categories in unlabeled data closer to these prototypes to form more compact clusters and avoid boundary overlap between known and novel categories. On the other hand, we use these prototypes to calibrate noisy prototypes estimated from unlabeled data based on category similarities, which allows for more accurate estimation of prototypes for novel categories that can be used as reliable learning targets later. After knowledge transfer, we further propose two feature alignment mechanisms to acquire both instance- and category-level knowledge from unlabeled data by aligning instance features with both augmented features and the calibrated prototypes, which can boost model performance on both known and novel categories with less noise. Experiments on three benchmark datasets show that our model outperforms SOTA methods, especially on novel categories. Theoretical analysis is provided for an in-depth understanding of our model in general. Our code and data are available at https://github.com/Lackel/TAN. Wenbin An, Feng Tian 0002, Wenkai Shi, Yan Chen 0031, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 1 |
| 2024 | A Unified Knowledge Transfer Network for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) aims to recognize both known and novel categories in an unlabeled dataset by leveraging another labeled dataset with only known categories. Without considering knowledge transfer from known to novel categories, current methods usually perform poorly on novel categories due to the lack of corresponding supervision. To mitigate this issue, we propose a unified Knowledge Transfer Network (KTN), which solves two obstacles to knowledge transfer in GCD. First, the mixture of known and novel categories in unlabeled data makes it difficult to identify transfer candidates (i.e., samples with novel categories). For this, we propose an entropy-based method that leverages knowledge in the pre-trained classifier to differentiate known and novel categories without requiring extra data or parameters. Second, the lack of prior knowledge of novel categories presents challenges in quantifying semantic relationships between categories to decide the transfer weights. For this, we model different categories with prototypes and treat their similarities as transfer weights to measure the semantic similarities between categories. On the basis of two treatments, we transfer knowledge from known to novel categories by conducting pre-adjustment of logits and post-adjustment of labels for transfer candidates based on the transfer weights between different categories. With the weighted adjustment, KTN can generate more accurate pseudo-labels for unlabeled data, which helps to learn more discriminative features and boost model performance on novel categories. Extensive experiments show that our method outperforms state-of-the-art models on all evaluation metrics across multiple benchmark datasets. Furthermore, different from previous clustering-based methods that can only work offline with abundant data, KTN can be deployed online conveniently with faster inference speed. Code and data are available at https://github.com/yibai-shi/KTN. Wenkai Shi, Wenbin An, Feng Tian 0002, Yan Chen 0031, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 2 |
| 2024 | A Tri-Branch Network with Prototype-aware Matching for Universal Category DiscoveryabstractIn this paper, we propose a novel task, Universal Category Discovery (UCD), to address the challenge of partial overlap between source and target domain categories. Different from previous tasks that assume all known categories exist in the target domain, UCD introduces "private-known" categories that only exist in the source domain and aims to classify unlabeled data as "common" or "novel" categories while avoiding misclassifying them into "private-known" categories. For this task, we propose a Tri-branch network with bidirectional Prototype-aware Matching (TriPM). TriPM effectively transfers knowledge from labeled to unlabeled data by bidirectionally matching similar data pairs, while a prototype matching strategy reduces the negative transfer risk from "private-known" categories. Finally, we propose a tri-branch network to decouple knowledge acquisition from labeled data, unlabeled data, and their interactions, which can avoid knowledge forgetting, explore novel patterns, and transfer common knowledge, respectively. Experiments demonstrate our model’s superiority over SOTA methods. Haonan Lin, Wenbin An, Yan Chen 0031, Feng Tian 0002, Yuzhe Yao, Wei Ding 0003, Qianying Wang 0002, Ping Chen 0001 |
ICME | 2 |
| 2024 | Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image EditingabstractText-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires inverting the source image into a latent space, a process often hindered by prediction errors inherent in DDIM inversion.
These errors accumulate during the diffusion process, resulting in inferior content preservation and edit fidelity, especially with conditional inputs.
We address these challenges by investigating the primary contributors to error accumulation in DDIM inversion and identify the singularity problem in traditional noise schedules as a key issue.
To resolve this, we introduce the *Logistic Schedule*, a novel noise schedule designed to eliminate singularities, improve inversion stability, and provide a better noise space for image editing. This schedule reduces noise prediction errors, enabling more faithful editing that preserves the original content of the source image. Our approach requires no additional retraining and is compatible with various existing editing methods.
Experiments across eight editing tasks demonstrate the Logistic Schedule's superior performance in content preservation and edit fidelity compared to traditional noise schedules, highlighting its adaptability and effectiveness.
The project page is available at https://lonelvino.github.io/SYE/. Haonan Lin, Yan Chen 0031, Jiahao Wang 0004, Wenbin An, Mengmeng Wang 0005, Feng Tian 0002, Yong Liu 0007, Guang Dai, Jingdong Wang 0001, Qianying Wang 0002 |
NeurIPS | 4 |
| 2024 | Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category DiscoveryabstractRecent advancements have shown promise in applying traditional Semi-Supervised Learning strategies to the task of Generalized Category Discovery (GCD). Typically, this involves a teacher-student framework in which the teacher imparts knowledge to the student to classify categories, even in the absence of explicit labels. Nevertheless, GCD presents unique challenges, particularly the absence of priors for new classes, which can lead to the teacher's misguidance and unsynchronized learning with the student, culminating in suboptimal outcomes. In our work, we delve into why traditional teacher-student designs falter in generalized category discovery as compared to their success in closed-world semi-supervised learning. We identify inconsistent pattern learning as the crux of this issue and introduce FlipClass—a method that dynamically updates the teacher to align with the student's attention, instead of maintaining a static teacher reference. Our teacher-attention-update strategy refines the teacher's focus based on student feedback, promoting consistent pattern recognition and synchronized learning across old and new classes. Extensive experiments on a spectrum of benchmarks affirm that FlipClass significantly surpasses contemporary GCD methods, establishing new standards for the field. Haonan Lin, Wenbin An, Jiahao Wang 0004, Yan Chen 0031, Feng Tian 0002, Mengmeng Wang 0005, Qianying Wang 0002, Guang Dai, Jingdong Wang 0001 |
NeurIPS | 2 |
| 2024 | DOWN: Dynamic Order Weighted Network for Fine-grained Category Discovery
Wenbin An, Feng Tian 0002, Wenkai Shi, Haonan Lin, Yaqiang Wu, Mingxiang Cai, Luyan Wang, Hua Wen, Ping Chen 0001 |
Knowl. Based Syst. | 1 |
| 2023 | Generalized Category Discovery with Decoupled Prototypical NetworkabstractGeneralized Category Discovery (GCD) aims to recognize both known and novel categories from a set of unlabeled data, based on another dataset labeled with only known categories. Without considering differences between known and novel categories, current methods learn about them in a coupled manner, which can hurt model's generalization and discriminative ability. Furthermore, the coupled training approach prevents these models transferring category-specific knowledge explicitly from labeled data to unlabeled data, which can lose high-level semantic information and impair model performance. To mitigate above limitations, we present a novel model called Decoupled Prototypical Network (DPN). By formulating a bipartite matching problem for category prototypes, DPN can not only decouple known and novel categories to achieve different training targets effectively, but also align known categories in labeled and unlabeled data to transfer category-specific knowledge explicitly and capture high-level semantics. Furthermore, DPN can learn more discriminative features for both known and novel categories through our proposed Semantic-aware Prototypical Learning (SPL). Besides capturing meaningful semantic information, SPL can also alleviate the noise of hard pseudo labels through semantic-weighted soft assignment. Extensive experiments show that DPN outperforms state-of-the-art models by a large margin on all evaluation metrics across multiple benchmark datasets. Code and data are available at https://github.com/Lackel/DPN. Wenbin An, Feng Tian 0002, Wei Ding 0003, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 1 |
| 2023 | DNA: Denoised Neighborhood Aggregation for Fine-grained Category DiscoveryabstractDiscovering fine-grained categories from coarsely labeled data is a practical and challenging task, which can bridge the gap between the demand for fine-grained analysis and the high annotation cost.Previous works mainly focus on instance-level discrimination to learn low-level features, but ignore semantic similarities between data, which may prevent these models learning compact cluster representations.In this paper, we propose Denoised Neighborhood Aggregation (DNA), a self-supervised framework that encodes semantic structures of data into the embedding space.Specifically, we retrieve k-nearest neighbors of a query as its positive keys to capture semantic similarities between data and then aggregate information from the neighbors to learn compact cluster representations, which can make fine-grained categories more separatable.However, the retrieved neighbors can be noisy and contain many false-positive keys, which can degrade the quality of learned embeddings.To cope with this challenge, we propose three principles to filter out these false neighbors for better representation learning.Furthermore, we theoretically justify that the learning objective of our framework is equivalent to a clustering loss, which can capture semantic similarities between data to form compact fine-grained clusters.Extensive experiments on three benchmark datasets show that our method can retrieve more accurate neighbors (21.31% accuracy improvement) and outperform state-of-the-art models by a large margin (average 9.96% improvement on three metrics).Our code and data are available at https://github.com/Lackel/DNA. Wenbin An, Feng Tian 0002, Wenkai Shi, Yan Chen 0031, Qianying Wang 0002, Ping Chen 0001 |
EMNLP | 1 |
| 2023 | A Diffusion Weighted Graph Framework for New Intent DiscoveryabstractNew Intent Discovery (NID) aims to recognize both new and known intents from unlabeled data with the aid of limited labeled data containing only known intents.Without considering structure relationships between samples, previous methods generate noisy supervisory signals which cannot strike a balance between quantity and quality, hindering the formation of new intent clusters and effective transfer of the pre-training knowledge.To mitigate this limitation, we propose a novel Diffusion Weighted Graph Framework (DWGF) to capture both semantic similarities and structure relationships inherent in data, enabling more sufficient and reliable supervisory signals.Specifically, for each sample, we diffuse neighborhood relationships along semantic paths guided by the nearest neighbors for multiple hops to characterize its local structure discriminately.Then, we sample its positive keys and weigh them based on semantic similarities and local structures for contrastive learning.During inference, we further propose Graph Smoothing Filter (GSF) to explicitly utilize the structure relationships to filter high-frequency noise embodied in semantically ambiguous samples on the cluster boundary.Extensive experiments show that our method outperforms state-of-the-art models on all evaluation metrics across multiple benchmark datasets. Wenkai Shi, Wenbin An, Feng Tian 0002, Qianying Wang 0002, Ping Chen 0001 |
EMNLP | 2 |
| 2023 | Aspect-Based Sentiment Analysis With Heterogeneous Graph Neural NetworkabstractAspect-based sentiment analysis aims to predict sentiment polarities of given aspects in text. Most current approaches employ attention-based neural methods to capture semantic relationships between aspects and words in one sentence. However, these methods ignore the fact that sentences with the same aspect and sentiment polarity often share the structure and semantic information in a domain, which leads to lower model performance. To mitigate this problem, we propose a heterogeneous aspect graph neural network (HAGNN) to learn the structure and semantic knowledge from intersentence relationships. Our model is a heterogeneous graph neural network since it contains three different kinds of nodes: word nodes, aspect nodes, and sentence nodes. These nodes can pass structure and semantic information between each other and update their embeddings to improve the performance of our model. To the best of our knowledge, we are the first to use a heterogeneous graph to capture relationships between sentences and aspects. The experimental results on five public datasets show the effectiveness of our model outperforming some state-of-the-art models. Wenbin An, Feng Tian 0002, Ping Chen 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2022 | Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive LearningabstractNovel category discovery aims at adapting models trained on known categories to novel categories.Previous works only focus on the scenario where known and novel categories are of the same granularity.In this paper, we investigate a new practical scenario called Fine-grained Category Discovery under Coarsegrained supervision (FCDC).FCDC aims at discovering fine-grained categories with only coarse-grained labeled data, which can adapt models to categories of different granularity from known ones and reduce significant labeling cost.It is also a challenging task since supervised training on coarse-grained categories tends to focus on inter-class distance (distance between coarse-grained classes) but ignore intra-class distance (distance between fine-grained sub-classes) which is essential for separating fine-grained categories.Considering most current methods cannot transfer knowledge from coarse-grained level to fine-grained level, we propose a hierarchical weighted self-contrastive network by building a novel weighted self-contrastive module and combining it with supervised learning in a hierarchical manner.Extensive experiments on public datasets show both effectiveness and efficiency of our model over compared methods.Code and data are available at https://github.com/Lackel/ Hierarchical_Weighted_SCL. Wenbin An, Feng Tian 0002, Ping Chen 0001, Siliang Tang, Qianying Wang 0002 |
EMNLP | 1 |