EDBT 2026 Demo / reviewers in the wild / expert
Wei Suo
dblp:251/2854
· DBLP profile ↗
23ranked-venue papers
10as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Look and check: A multi-label classification pipeline via multi-agent cooperation
Mingyu Fu, Wei Suo, Lingyan Ran, Peng Wang 0015 |
Neurocomputing | 2 |
| 2026 | Semi-Supervised VQA Multi-Modal Explanation via Self-Critical LearningabstractVQA explanation task aims to explain the decision-making process of VQA models in a way that is easily understandable to humans. Existing methods mostly use visual location or natural language explanation approaches to generate corresponding rationales. Although significant progress has been made, these frameworks are bottlenecked by the following challenges: 1) Uni-modal paradigm inevitably leads to semantic ambiguity of explanations. 2) The reasoning process cannot be faithfully responded to and suffers from logical inconsistency. 3) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we introduce a new Semi-supervised VQA Multi-modal Explanation (SME) method via self-critical learning, which addresses the above challenges by leveraging both visual and textual explanations to comprehensively reveal the inference process of the model. Meanwhile, in order to improve the logical consistency between answers and rationales, we design a novel self-critical strategy to evaluate candidate explanations based on answer reward scores. More importantly, our method can benefit from a tremendous amount of samples without human-annotated explanations with semi-supervised learning. Extensive automatic measures and human evaluations all show the effectiveness of our method. Finally, the framework achieves a new state-of-the-art performance on the three VQA explanation datasets. Wei Suo, Ji Ma 0008, Mengyang Sun, Hanwang Zhang, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | More video-relevant paragraph captioning via Perturbed Attention Self-Distillation
Yiqi Gao, Wei Suo, Mengyang Sun, Le Liu 0008, Peng Wang 0015 |
Pattern Recognit. | 2 |
| 2025 | Octopus: Alleviating Hallucination via Dynamic Contrastive DecodingabstractLarge Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from serious hallucination problems and tend to generate fabricated responses. Recently, several Contrastive Decoding (CD) strategies have been proposed to alleviate hallucination by introducing disturbed inputs. Although great progress has been made, these CD strategies mostly apply a one-size-fits-all approach for all input conditions. In this paper, we revisit this process through extensive experiments. Related results show that hallucination causes are hybrid and each generative step faces a unique hallucination challenge. Leveraging these meaningful insights, we introduce a simple yet effective Octopus-Like framework that enables the model to adaptively identify hallucination types and create a dynamic CD workflow. Our Octopus framework not only outperforms existing methods across four benchmarks but also demonstrates excellent deployability and expansibility. Code is available at https://github.com/LijunZhang01/Octopus. Wei Suo, Mengyang Sun, Lin Wu 0001, Peng Wang 0015 |
CVPR | 1 |
| 2025 | Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language ModelsabstractAlthough Large Vision-Language Models (LVLMs) have achieved impressive results, their high computational costs pose a significant barrier to wide application. To enhance inference efficiency, most existing approaches can be categorized as parameter-dependent or token-dependent strategies to reduce computational demands. However, parameter-dependent methods require retraining LVLMs to recover performance while token-dependent strategies struggle to consistently select the most relevant tokens. In this paper, we systematically analyze the above challenges and provide a series of valuable insights for inference acceleration. Based on these findings, we propose a novel framework, the Pruning All-Rounder (PAR). Different from previous works, PAR develops a meta-router to adaptively organize pruning flows across both tokens and layers. With a self-supervised learning manner, our method achieves a superior balance between performance and efficiency. Notably, PAR is highly flexible, offering multiple pruning versions to address a range of acceleration scenarios. The code for this work is publicly available at https://github.com/ASGO-MM/Pruning-All-Rounder. Wei Suo, Ji Ma 0008, Mengyang Sun, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001 |
ICCV | 1 |
| 2025 | Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
Ji Ma 0008, Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 2 |
| 2025 | Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
Mingyu Fu, Wei Suo, Ji Ma 0008, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 2 |
| 2025 | CHA: Conditional Hyper-Adapter method for detecting human-object interaction
Mengyang Sun, Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2025 | Distributed online constrained nonconvex optimization in dynamic environments over directed graphs
Wei Suo, Wenling Li, Yang Liu 0096, Jia Song 0002 |
Signal Process. | 1 |
| 2025 | CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary LearningabstractThis article investigates the problem of continual learning (CL) of vision-language models (VLMs) in open domains, where models are required to perform continual updating and inference on a stream of datasets from diverse seen and unseen domains with novel classes. Such a capability is crucial for various applications in open environments, e.g., AI assistants, autonomous driving systems, and robotics. Current CL studies mostly focus on closed-set scenarios in a single domain with known classes. Large pretrained VLMs such as CLIP have showcased exceptional zero-shot recognition capabilities, and several recent studies have leveraged the unique characteristics of VLMs to mitigate catastrophic forgetting in CL. However, they primarily focus on closed-set CL in a single-domain dataset. Open-domain CL of large VLMs is significantly more challenging due to 1) large class correlations and domain gaps across the datasets and 2) the forgetting of zero-shot knowledge in the pretrained VLMs and the knowledge learned from the newly adapted datasets. In this work, we introduce a novel approach, termed CoLeCLIP, which learns an open-domain CL model based on CLIP. It addresses these challenges through joint learning of a set of task prompts and a cross-domain class vocabulary. Extensive experiments on 11 domain datasets show that CoLeCLIP achieves new state-of-the-art performance for open-domain CL under both task- and class-incremental learning (CIL) settings. Guansong Pang, Wei Suo, Chenchen Jing, Yuling Xi, Lingqiao Liu, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Distributed Online Convex Optimization Over Time-Varying Unbalanced Digraphs With Multiple Coupled ConstraintsabstractThis article aims to solve the distributed online convex optimization (DOCO) problems subjected to multiple coupled constraints over time-varying (TV) unbalanced digraphs. The existing global constraint models and coupled constraint models, where the number of constraints is related to the number of nodes, are not sufficient to reflect the characteristics of multiple coupled constrained optimization problems. On account of this drawback, a multiple coupled constraint model is constructed, which contains several coupled constraints, only including a part of all nodes. In addition, practical TV scenarios commonly come with complex network connectivity, which requires diverse matrices for fusing various information. In view of connectivity requirements, a novel TV distributed primal–dual push–pull (TDPP) algorithm, which can convert two types of weight matrices to all onefold row stochastic (RS) matrices, is proposed to tackle multiple coupled constrained problems. Under some general and necessary assumptions and conditions, both desired sublinear dynamic regret and constraint violation can be acquired by a strict theoretical analysis. Finally, two numerical examples are utilized to verify the superiority and validity of the TDPP algorithm compared with similar algorithms. Wei Suo, Wenling Li, Bin Zhang 0023, Yang Liu 0096 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Rethinking and Improving Visual Prompt Selection for In-Context Learning Segmentation
Wei Suo, Lanqing Lai, Mengyang Sun, Hanwang Zhang, Peng Wang 0015, Yanning Zhang 0001 |
ECCV (46) | 1 |
| 2024 | C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning
Ji Ma 0008, Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
IJCAI | 2 |
| 2024 | A Plug-and-Play Method for Rare Human-Object Interactions Detection by Bridging Domain GapabstractHuman-object interactions (HOI) detection aims at capturing human-object pairs in images and corresponding actions. It is an important step toward high-level visual reasoning and scene understanding. However, due to the natural bias from the real world, existing methods mostly struggle with rare human-object pairs and lead to sub-optimal results. Recently, with the development of the generative model, a straightforward approach is to construct a more balanced dataset based on a group of supplementary samples. Unfortunately, there is a significant domain gap between the generated data and the original data, and simply merging the generated images into the original dataset cannot significantly boost the performance. To alleviate the above problem, we present a novel model-agnostic framework called Context-Enhanced Feature Alignment (CEFA) module, which can effectively align the generated data with the original data at the feature level and bridge the domain gap. Specifically, CEFA consists of a feature alignment module and a context enhancement module. On one hand, considering the crucial role of human-object pairs information in HOI tasks, the feature alignment module aligns the human-object pairs by aggregating instance information. On the other hand, to mitigate the issue of losing important context information caused by the traditional discriminator-style alignment method, we employ a context-enhanced image reconstruction module to improve the model's learning ability of contextual cues. Extensive experiments have shown that our method can serve as a plug-and-play module to improve the detection performance of HOI models on rare categories. Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 2 |
| 2024 | An Adaptive Correlation Filtering Method for Text-Based Person Search
Mengyang Sun, Wei Suo, Peng Wang 0015, Kai Niu 0002, Le Liu 0008, Guosheng Lin, Yanning Zhang 0001, Qi Wu 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical LearningabstractVQA Natural Language Explanation (VQA-NLE) task aims to explain the decision-making process of VQA models in natural language. Unlike traditional attention or gradient analysis, free-text rationales can be easier to understand and gain users' trust. Existing methods mostly use post-hoc or selfrationalization models to obtain a plausible explanation. However, these frameworks are bottle-necked by the following challenges: 1) the reasoning process cannot be faithfully responded to and suffer from the problem of logical inconsistency. 2) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we propose a new Semi-Supervised VQA-NLE via Self-Critical Learning (S3C), which evaluates the candidate explanations by answering rewards to improve the logical consistency between answers and rationales. With a semi-supervised learning framework, the S3C can benefit from a tremendous amount of samples without human-annotated explanations. A large number of automatic measures and human evaluations all show the effectiveness of our method. Meanwhile, the framework achieves a new state-of-the-art performance on the two VQA-NLE datasets. Wei Suo, Mengyang Sun, Weisong Liu, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
CVPR | 1 |
| 2023 | AHT: A Novel Aggregation Hyper-transformer for Few-Shot Object Detection
Lanqing Lai, Yale Yu, Wei Suo, Peng Wang 0015 |
PRCV (12) | 3 |
| 2023 | Rethinking and Improving Feature Pyramids for One-Stage Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) is an important task in the vision-and-language community, since it is an essential step for many cross-modal tasks such as VQA, image retrieval and image caption. To obtain a better trade-off between speed and accuracy, existing researches usually follow a one-stage paradigm, where this task can be considered as a language-conditioned object detection task. Meanwhile, previous one-stage REC frameworks provide many different research perspectives, such as the strategies of fusion, the stage of fusion and the design of detection head. Surprisingly, these works mostly ignore the value of integrating multi-level features and even only apply single-scale features to locate the target. In this paper, we focus on rethinking and improving feature pyramids for one-stage REC. By experimental validations, we first prove that although multi-scale fusion is an effective approach for improving performance, the mature neck structures from object detection (e.g., FPN, BFN and HRFPN) have a limited impact on this task. Further, we visualize the outputs of FPN and find the underlying reason is that these coarse-grained FPN fusion strategies suffer from semantic ambiguity problem. Based on the above insights, we propose a new Language-Guided FPN (LG-FPN) method, which can dynamically allocate and select the fine-grained information by stacking language-gate and union-gate. A large number of contrastive and ablative experiments show that our LG-FPN is an effective and reliable module that can adapt to different visual backbones, fusion strategies and detection heads. Finally, our method achieves state-of-the-art performance on four referring expression datasets. Wei Suo, Mengyang Sun, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | A Proposal-Free One-Stage Framework for Referring Expression Comprehension and Generation via Dense Cross-AttentionabstractReferring Expression Comprehension (REC) and Generation (REG) have become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering or visual dialogue. However, it has not been widely used in many downstream tasks, mainly for the following reasons: 1) mainstream two-stage methods rely on additional annotations or off-the-shelf detectors to generate proposals. It would heavily degrade the generalization ability of models and lead to inevitable error accumulation. 2) Although one-stage strategies for REC have been proposed, these methods have to depend on lots of hyper-parameters (such as anchors) to generate bounding box. In this paper, we present a proposal-free one-stage (PFOS) framework that can directly regress the region-of-interest from the image or generate unambiguous descriptions in an end-to-end manner. Instead of using the dominant two-stage fashion, we take the dense-grid of images as input for a cross-attention transformer that learns multi-modal correspondences. The final bounding box or sentence is directly predicted from the image without the anchor selection or the computation of visual difference. Furthermore, we expand the traditional two-stage listener-speaker framework to jointly train by a one-stage learning paradigm. Our model achieves state-of-the-art performance on both accuracy and speed for comprehension and competitive results for generation. Mengyang Sun, Wei Suo, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | A Simple and Robust Correlation Filtering Method for Text-Based Person Search
Wei Suo, Mengyang Sun, Kai Niu 0002, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
ECCV (35) | 1 |
| 2022 | Dual-Level Decoupled Transformer for Video CaptioningabstractVideo captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generator to learn from offline-extracted motion or appearance features from pre-trained vision models. However, these methods may suffer from the so-called "couple" drawbacks on both video spatio-temporal representation and sentence generation. For the former, "couple" means learning spatio-temporal representation in a single model(3DCNN), resulting the problems named disconnection in task/pre-train domain and hard for end-to-end training. As for the latter, "couple" means treating the generation of visual semantic and syntax-related words equally. To this end, we present D2 - a dual-level decoupled transformer pipeline to solve the above drawbacks: (i) for video spatio-temporal representation, we decouple the process of it into "first-spatial-then-temporal" paradigm, releasing the potential of using dedicated model(e.g. image-text pre-training) to connect the pre-training and downstream tasks, and makes the entire model end-to-end trainable. (ii) for sentence generation, we propose Syntax-Aware Decoder to dynamically measure the contribution of visual semantic and syntax-related words. Extensive experiments on three widely-used benchmarks (MSVD, MSR-VTT and VATEX) have shown great potential of the proposed D2 and surpassed the previous methods by a large margin in the task of video captioning. Yiqi Gao, Xinglin Hou, Wei Suo, Mengyang Sun, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015 |
ICMR | 3 |
| 2022 | Improving Image Captioning via Enhancing Dual-Side Context AwarenessabstractRecent work on visual question answering demonstrate that grid features can work as well as region feature on vision language tasks. In the meantime, transformer-based model and its variants have shown remarkable performance on image captioning. However, the object-contextual information missing caused by the single granularity nature of grid feature on the encoder side, as well as the future contextual information missing due to the left2right decoding paradigm of transformer decoder, remains unexplored. In this work, we tackle these two problems by enhancing contextual information at dual-side:(i) at encoder side, we propose Context-Aware Self-Attention module, in which the key/value is expanded with adjacent rectangle region where each region contains two or more aggregated grid features; this enables grid feature with varying granularity, storing adequate contextual information for object with different scale. (ii) at decoder side, we incorporate a dual-way decoding strategy, in which left2right and right2left decoding are conducted simultaneously and interactively. It utilizes both past and future contextual information when generates current word. Combining these two modules with a vanilla transformer, our Context-Aware Transformer(CATNet) achieves a new state-of-the-art on MSCOCO benchmark. Yiqi Gao, Ning Wang 0020, Wei Suo, Mengyang Sun, Peng Wang 0015 |
ICMR | 3 |
| 2021 | Proposal-free One-stage Referring Expression via Grid-Word Cross-AttentionabstractReferring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering. However, it has not been widely used in many downstream tasks because it suffers 1) two-stage methods exist heavy computation cost and inevitable error accumulation, and 2) one-stage methods have to depend on lots of hyper-parameters (such as anchors) to generate bounding box. In this paper, we present a proposal-free one-stage (PFOS) model that is able to regress the region-of-interest from the image, based on a textual query, in an end-to-end manner. Instead of using the dominant anchor proposal fashion, we directly take the dense-grid of image as input for a cross-attention transformer that learns grid-word correspondences. The final bounding box is predicted directly from the image without the time-consuming anchor selection process that previous methods suffer. Our model achieves the state-of-the-art performance on four referring expression datasets with higher efficiency, comparing to previous best one-stage and two-stage methods. Wei Suo, Mengyang Sun, Peng Wang 0015, Qi Wu 0001 |
IJCAI | 1 |