Wei Suo

dblp:251/2854 · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Look and check: A multi-label classification pipeline via multi-agent cooperation
Mingyu Fu, Wei Suo, Lingyan Ran, Peng Wang 0015
Neurocomputing2
2026 Semi-Supervised VQA Multi-Modal Explanation via Self-Critical Learning
abstract
VQA explanation task aims to explain the decision-making process of VQA models in a way that is easily understandable to humans. Existing methods mostly use visual location or natural language explanation approaches to generate corresponding rationales. Although significant progress has been made, these frameworks are bottlenecked by the following challenges: 1) Uni-modal paradigm inevitably leads to semantic ambiguity of explanations. 2) The reasoning process cannot be faithfully responded to and suffers from logical inconsistency. 3) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we introduce a new Semi-supervised VQA Multi-modal Explanation (SME) method via self-critical learning, which addresses the above challenges by leveraging both visual and textual explanations to comprehensively reveal the inference process of the model. Meanwhile, in order to improve the logical consistency between answers and rationales, we design a novel self-critical strategy to evaluate candidate explanations based on answer reward scores. More importantly, our method can benefit from a tremendous amount of samples without human-annotated explanations with semi-supervised learning. Extensive automatic measures and human evaluations all show the effectiveness of our method. Finally, the framework achieves a new state-of-the-art performance on the three VQA explanation datasets.
Wei Suo, Ji Ma 0008, Mengyang Sun, Hanwang Zhang, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 More video-relevant paragraph captioning via Perturbed Attention Self-Distillation
Yiqi Gao, Wei Suo, Mengyang Sun, Le Liu 0008, Peng Wang 0015
Pattern Recognit.2
2025 Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding
abstract
Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from serious hallucination problems and tend to generate fabricated responses. Recently, several Contrastive Decoding (CD) strategies have been proposed to alleviate hallucination by introducing disturbed inputs. Although great progress has been made, these CD strategies mostly apply a one-size-fits-all approach for all input conditions. In this paper, we revisit this process through extensive experiments. Related results show that hallucination causes are hybrid and each generative step faces a unique hallucination challenge. Leveraging these meaningful insights, we introduce a simple yet effective Octopus-Like framework that enables the model to adaptively identify hallucination types and create a dynamic CD workflow. Our Octopus framework not only outperforms existing methods across four benchmarks but also demonstrates excellent deployability and expansibility. Code is available at https://github.com/LijunZhang01/Octopus.
Wei Suo, Mengyang Sun, Lin Wu 0001, Peng Wang 0015
CVPR1
2025 Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models
abstract
Although Large Vision-Language Models (LVLMs) have achieved impressive results, their high computational costs pose a significant barrier to wide application. To enhance inference efficiency, most existing approaches can be categorized as parameter-dependent or token-dependent strategies to reduce computational demands. However, parameter-dependent methods require retraining LVLMs to recover performance while token-dependent strategies struggle to consistently select the most relevant tokens. In this paper, we systematically analyze the above challenges and provide a series of valuable insights for inference acceleration. Based on these findings, we propose a novel framework, the Pruning All-Rounder (PAR). Different from previous works, PAR develops a meta-router to adaptively organize pruning flows across both tokens and layers. With a self-supervised learning manner, our method achieves a superior balance between performance and efficiency. Notably, PAR is highly flexible, offering multiple pruning versions to address a range of acceleration scenarios. The code for this work is publicly available at https://github.com/ASGO-MM/Pruning-All-Rounder.
Wei Suo, Ji Ma 0008, Mengyang Sun, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001
ICCV1
2025 Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
Ji Ma 0008, Wei Suo, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia2
2025 Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
Mingyu Fu, Wei Suo, Ji Ma 0008, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia2
2025 CHA: Conditional Hyper-Adapter method for detecting human-object interaction
Mengyang Sun, Wei Suo, Peng Wang 0015, Yanning Zhang 0001
Pattern Recognit.2
2025 Distributed online constrained nonconvex optimization in dynamic environments over directed graphs
Wei Suo, Wenling Li, Yang Liu 0096, Jia Song 0002
Signal Process.1
2025 CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary Learning
abstract
This article investigates the problem of continual learning (CL) of vision-language models (VLMs) in open domains, where models are required to perform continual updating and inference on a stream of datasets from diverse seen and unseen domains with novel classes. Such a capability is crucial for various applications in open environments, e.g., AI assistants, autonomous driving systems, and robotics. Current CL studies mostly focus on closed-set scenarios in a single domain with known classes. Large pretrained VLMs such as CLIP have showcased exceptional zero-shot recognition capabilities, and several recent studies have leveraged the unique characteristics of VLMs to mitigate catastrophic forgetting in CL. However, they primarily focus on closed-set CL in a single-domain dataset. Open-domain CL of large VLMs is significantly more challenging due to 1) large class correlations and domain gaps across the datasets and 2) the forgetting of zero-shot knowledge in the pretrained VLMs and the knowledge learned from the newly adapted datasets. In this work, we introduce a novel approach, termed CoLeCLIP, which learns an open-domain CL model based on CLIP. It addresses these challenges through joint learning of a set of task prompts and a cross-domain class vocabulary. Extensive experiments on 11 domain datasets show that CoLeCLIP achieves new state-of-the-art performance for open-domain CL under both task- and class-incremental learning (CIL) settings.
Guansong Pang, Wei Suo, Chenchen Jing, Yuling Xi, Lingqiao Liu, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015
IEEE Trans. Neural Networks Learn. Syst.3
2025 Distributed Online Convex Optimization Over Time-Varying Unbalanced Digraphs With Multiple Coupled Constraints
abstract
This article aims to solve the distributed online convex optimization (DOCO) problems subjected to multiple coupled constraints over time-varying (TV) unbalanced digraphs. The existing global constraint models and coupled constraint models, where the number of constraints is related to the number of nodes, are not sufficient to reflect the characteristics of multiple coupled constrained optimization problems. On account of this drawback, a multiple coupled constraint model is constructed, which contains several coupled constraints, only including a part of all nodes. In addition, practical TV scenarios commonly come with complex network connectivity, which requires diverse matrices for fusing various information. In view of connectivity requirements, a novel TV distributed primal–dual push–pull (TDPP) algorithm, which can convert two types of weight matrices to all onefold row stochastic (RS) matrices, is proposed to tackle multiple coupled constrained problems. Under some general and necessary assumptions and conditions, both desired sublinear dynamic regret and constraint violation can be acquired by a strict theoretical analysis. Finally, two numerical examples are utilized to verify the superiority and validity of the TDPP algorithm compared with similar algorithms.
Wei Suo, Wenling Li, Bin Zhang 0023, Yang Liu 0096
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Rethinking and Improving Visual Prompt Selection for In-Context Learning Segmentation
Wei Suo, Lanqing Lai, Mengyang Sun, Hanwang Zhang, Peng Wang 0015, Yanning Zhang 0001
ECCV (46)1
2024 C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning
Ji Ma 0008, Wei Suo, Peng Wang 0015, Yanning Zhang 0001
IJCAI2
2024 A Plug-and-Play Method for Rare Human-Object Interactions Detection by Bridging Domain Gap
abstract
Human-object interactions (HOI) detection aims at capturing human-object pairs in images and corresponding actions. It is an important step toward high-level visual reasoning and scene understanding. However, due to the natural bias from the real world, existing methods mostly struggle with rare human-object pairs and lead to sub-optimal results. Recently, with the development of the generative model, a straightforward approach is to construct a more balanced dataset based on a group of supplementary samples. Unfortunately, there is a significant domain gap between the generated data and the original data, and simply merging the generated images into the original dataset cannot significantly boost the performance. To alleviate the above problem, we present a novel model-agnostic framework called Context-Enhanced Feature Alignment (CEFA) module, which can effectively align the generated data with the original data at the feature level and bridge the domain gap. Specifically, CEFA consists of a feature alignment module and a context enhancement module. On one hand, considering the crucial role of human-object pairs information in HOI tasks, the feature alignment module aligns the human-object pairs by aggregating instance information. On the other hand, to mitigate the issue of losing important context information caused by the traditional discriminator-style alignment method, we employ a context-enhanced image reconstruction module to improve the model's learning ability of contextual cues. Extensive experiments have shown that our method can serve as a plug-and-play module to improve the detection performance of HOI models on rare categories.
Wei Suo, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia2
2024 An Adaptive Correlation Filtering Method for Text-Based Person Search
Mengyang Sun, Wei Suo, Peng Wang 0015, Kai Niu 0002, Le Liu 0008, Guosheng Lin, Yanning Zhang 0001, Qi Wu 0001
Int. J. Comput. Vis.2
2023 S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical Learning
abstract
VQA Natural Language Explanation (VQA-NLE) task aims to explain the decision-making process of VQA models in natural language. Unlike traditional attention or gradient analysis, free-text rationales can be easier to understand and gain users' trust. Existing methods mostly use post-hoc or selfrationalization models to obtain a plausible explanation. However, these frameworks are bottle-necked by the following challenges: 1) the reasoning process cannot be faithfully responded to and suffer from the problem of logical inconsistency. 2) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we propose a new Semi-Supervised VQA-NLE via Self-Critical Learning (S3C), which evaluates the candidate explanations by answering rewards to improve the logical consistency between answers and rationales. With a semi-supervised learning framework, the S3C can benefit from a tremendous amount of samples without human-annotated explanations. A large number of automatic measures and human evaluations all show the effectiveness of our method. Meanwhile, the framework achieves a new state-of-the-art performance on the two VQA-NLE datasets.
Wei Suo, Mengyang Sun, Weisong Liu, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
CVPR1
2023 AHT: A Novel Aggregation Hyper-transformer for Few-Shot Object Detection
Lanqing Lai, Yale Yu, Wei Suo, Peng Wang 0015
PRCV (12)3
2023 Rethinking and Improving Feature Pyramids for One-Stage Referring Expression Comprehension
abstract
Referring Expression Comprehension (REC) is an important task in the vision-and-language community, since it is an essential step for many cross-modal tasks such as VQA, image retrieval and image caption. To obtain a better trade-off between speed and accuracy, existing researches usually follow a one-stage paradigm, where this task can be considered as a language-conditioned object detection task. Meanwhile, previous one-stage REC frameworks provide many different research perspectives, such as the strategies of fusion, the stage of fusion and the design of detection head. Surprisingly, these works mostly ignore the value of integrating multi-level features and even only apply single-scale features to locate the target. In this paper, we focus on rethinking and improving feature pyramids for one-stage REC. By experimental validations, we first prove that although multi-scale fusion is an effective approach for improving performance, the mature neck structures from object detection (e.g., FPN, BFN and HRFPN) have a limited impact on this task. Further, we visualize the outputs of FPN and find the underlying reason is that these coarse-grained FPN fusion strategies suffer from semantic ambiguity problem. Based on the above insights, we propose a new Language-Guided FPN (LG-FPN) method, which can dynamically allocate and select the fine-grained information by stacking language-gate and union-gate. A large number of contrastive and ablative experiments show that our LG-FPN is an effective and reliable module that can adapt to different visual backbones, fusion strategies and detection heads. Finally, our method achieves state-of-the-art performance on four referring expression datasets.
Wei Suo, Mengyang Sun, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
IEEE Trans. Image Process.1
2023 A Proposal-Free One-Stage Framework for Referring Expression Comprehension and Generation via Dense Cross-Attention
abstract
Referring Expression Comprehension (REC) and Generation (REG) have become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering or visual dialogue. However, it has not been widely used in many downstream tasks, mainly for the following reasons: 1) mainstream two-stage methods rely on additional annotations or off-the-shelf detectors to generate proposals. It would heavily degrade the generalization ability of models and lead to inevitable error accumulation. 2) Although one-stage strategies for REC have been proposed, these methods have to depend on lots of hyper-parameters (such as anchors) to generate bounding box. In this paper, we present a proposal-free one-stage (PFOS) framework that can directly regress the region-of-interest from the image or generate unambiguous descriptions in an end-to-end manner. Instead of using the dominant two-stage fashion, we take the dense-grid of images as input for a cross-attention transformer that learns multi-modal correspondences. The final bounding box or sentence is directly predicted from the image without the anchor selection or the computation of visual difference. Furthermore, we expand the traditional two-stage listener-speaker framework to jointly train by a one-stage learning paradigm. Our model achieves state-of-the-art performance on both accuracy and speed for comprehension and competitive results for generation.
Mengyang Sun, Wei Suo, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
IEEE Trans. Multim.2
2022 A Simple and Robust Correlation Filtering Method for Text-Based Person Search
Wei Suo, Mengyang Sun, Kai Niu 0002, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
ECCV (35)1
2022 Dual-Level Decoupled Transformer for Video Captioning
abstract
Video captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generator to learn from offline-extracted motion or appearance features from pre-trained vision models. However, these methods may suffer from the so-called "couple" drawbacks on both video spatio-temporal representation and sentence generation. For the former, "couple" means learning spatio-temporal representation in a single model(3DCNN), resulting the problems named disconnection in task/pre-train domain and hard for end-to-end training. As for the latter, "couple" means treating the generation of visual semantic and syntax-related words equally. To this end, we present D2 - a dual-level decoupled transformer pipeline to solve the above drawbacks: (i) for video spatio-temporal representation, we decouple the process of it into "first-spatial-then-temporal" paradigm, releasing the potential of using dedicated model(e.g. image-text pre-training) to connect the pre-training and downstream tasks, and makes the entire model end-to-end trainable. (ii) for sentence generation, we propose Syntax-Aware Decoder to dynamically measure the contribution of visual semantic and syntax-related words. Extensive experiments on three widely-used benchmarks (MSVD, MSR-VTT and VATEX) have shown great potential of the proposed D2 and surpassed the previous methods by a large margin in the task of video captioning.
Yiqi Gao, Xinglin Hou, Wei Suo, Mengyang Sun, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015
ICMR3
2022 Improving Image Captioning via Enhancing Dual-Side Context Awareness
abstract
Recent work on visual question answering demonstrate that grid features can work as well as region feature on vision language tasks. In the meantime, transformer-based model and its variants have shown remarkable performance on image captioning. However, the object-contextual information missing caused by the single granularity nature of grid feature on the encoder side, as well as the future contextual information missing due to the left2right decoding paradigm of transformer decoder, remains unexplored. In this work, we tackle these two problems by enhancing contextual information at dual-side:(i) at encoder side, we propose Context-Aware Self-Attention module, in which the key/value is expanded with adjacent rectangle region where each region contains two or more aggregated grid features; this enables grid feature with varying granularity, storing adequate contextual information for object with different scale. (ii) at decoder side, we incorporate a dual-way decoding strategy, in which left2right and right2left decoding are conducted simultaneously and interactively. It utilizes both past and future contextual information when generates current word. Combining these two modules with a vanilla transformer, our Context-Aware Transformer(CATNet) achieves a new state-of-the-art on MSCOCO benchmark.
Yiqi Gao, Ning Wang 0020, Wei Suo, Mengyang Sun, Peng Wang 0015
ICMR3
2021 Proposal-free One-stage Referring Expression via Grid-Word Cross-Attention
abstract
Referring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering. However, it has not been widely used in many downstream tasks because it suffers 1) two-stage methods exist heavy computation cost and inevitable error accumulation, and 2) one-stage methods have to depend on lots of hyper-parameters (such as anchors) to generate bounding box. In this paper, we present a proposal-free one-stage (PFOS) model that is able to regress the region-of-interest from the image, based on a textual query, in an end-to-end manner. Instead of using the dominant anchor proposal fashion, we directly take the dense-grid of image as input for a cross-attention transformer that learns grid-word correspondences. The final bounding box is predicted directly from the image without the time-consuming anchor selection process that previous methods suffer. Our model achieves the state-of-the-art performance on four referring expression datasets with higher efficiency, comparing to previous best one-stage and two-stage methods.
Wei Suo, Mengyang Sun, Peng Wang 0015, Qi Wu 0001
IJCAI1