EDBT 2026 Demo / reviewers in the wild / expert
Wangbo Zhao
dblp:289/6986
· DBLP profile ↗
27ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0001-9545-7991ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education. Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang |
AAAI | 7 |
| 2026 | GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMsabstractWeidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Can Zhang, Xinyan Wan, Zhiyuan Liang, Pengfei Zhou, Yang You, Wangbo Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Xinyan Wan, Zhiyuan Liang, Yang You 0001, Wangbo Zhao |
ACL (1) | 10 |
| 2026 | Noisy correspondence decomposition for robust composed image retrieval
Zhaopan Xu, Chenrui Zhou, Xi Chen 0110, Wangbo Zhao, Jianning Zhang, Xiaojiang Peng, Hongxun Yao, Kaipeng Zhang |
Expert Syst. Appl. | 5 |
| 2026 | DyDiT++: Diffusion Transformers With Timestep and Spatial Dynamics for Efficient Visual GenerationabstractDiffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs primarily stem from the static inference paradigm, which inevitably introduces redundant computation in certain diffusion timesteps and spatial regions. To overcome this inefficiency, we propose Dynamic Diffusion Transformer (DyDiT), an architecture that dynamically adjusts its computation along both timestep and spatial dimensions. Specifically, we introduce a Timestep-wise Dynamic Width (TDW) approach that adapts model width conditioned on the generation timesteps. In addition, we design a Spatial-wise Dynamic Token (SDT) strategy to avoid redundant computation at unnecessary spatial locations. TDW and SDT can be seamlessly integrated into DiT and significantly accelerate the generation process. Building on these designs, we present an extended version, DyDiT++, with improvements in three key aspects. First, it extends the generation mechanism of DyDiT beyond diffusion to flow matching, demonstrating that our method can also accelerate flow-matching-based generation, enhancing its versatility. Furthermore, we enhance DyDiT to tackle more complex visual generation tasks, including video generation and text-to-image generation, thereby broadening its real-world applications. Finally, to address the high cost of full fine-tuning and democratize technology access, we investigate the feasibility of training DyDiT in a parameter-efficient manner and introduce timestep-based dynamic LoRA (TD-LoRA). Extensive experiments on diverse visual generation models, including DiT, SiT, Latte, and FLUX, demonstrate the effectiveness of DyDiT++. Remarkably, with $< $<3% additional fine-tuning iterations, our approach reduces the FLOPs of DiT-XL by 51%, yielding 1.73× realistic speedup on hardware, and achieves a competitive FID score of 2.07 on ImageNet. Wangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang 0036, Hao Luo 0004, Yibing Song, Gao Huang 0001, Fan Wang 0019, Yang You 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Training-Free Noisy Correspondence Rectification With Multimodal Conceptual Knowledge
Zhaopan Xu, Wangbo Zhao, Xi Chen 0110, Sheng Jin 0002, Hongxun Yao |
IEEE Signal Process. Lett. | 2 |
| 2026 | Learning With Dual Noisy Labels for Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification (TIReID) aims to identify a target person from a given textual description. Although recent work has made significant progress, most of it implicitly assumes that the sample annotations are correct and that the cross-modal correspondence in each image-text pair is well aligned. However, such an assumption requires elaborately annotated datasets, which are expensive and even impossible to obtain in practice. To alleviate this issue, in this letter, we explore a new TIReID setting, termed learning with dual noisy labels, in which the model learns from data with both noisy identity labels and noisy correspondence. We propose a general framework called TDTD (Two stage framework forDual noise ofTIReID) to achieve this. In the first stage, a Noise-Aware Preliminary Learning (NAPL) strategy selects “easy” triplets to train a noise-tolerant initial model. In the second stage, the model leverages reliable representations from NAPL to automatically correct both identity and correspondence errors via soft-label estimation and is then fine-tuned on the entire dataset using a dual noise-robust triplet loss. Extensive experiments on three public benchmarks, CUHK-PEDES, ICFG-PEDES, and RSTPReID, demonstrate the performance and robustness of TDTD, achieving state-of-the-art results under dual noise conditions. Zhaopan Xu, Wangbo Zhao, Xiaojiang Peng, Hongxun Yao |
IEEE Signal Process. Lett. | 2 |
| 2025 | A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMsabstractVision-language models (VLMs) have shown remarkable success across various multi-modal tasks, yet large VLMs encounter significant efficiency challenges due to processing numerous visual tokens. A promising approach to accelerating large VLM inference is using partial information, such as attention maps from specific layers, to assess token importance and prune less essential tokens. However, our study reveals three key insights: (i) Partial attention information is insufficient for accurately identifying critical visual tokens, resulting in suboptimal performance, especially at low token retention ratios; (ii) Global attention information, such as the attention map aggregated across all layers, more effectively preserves essential tokens and maintains comparable performance under aggressive pruning. However, the attention maps from all layers require a full inference pass, which increases computational load and is therefore impractical in existing methods; and (iii) The global attention map aggregated from a small VLM closely resembles that of a large VLM, suggesting an efficient alternative. Based on these findings, we introduce a training-free method, Small VLM Guidance for accelerating Large VLMs (SGL). Specifically, we employ the attention map aggregated from a small VLM to guide visual token pruning in a large VLM. Additionally, an early exiting mechanism is developed to fully use the small VLM’s predictions, dynamically invoking the larger VLM only when necessary, yielding a superior trade-off between accuracy and computation. Extensive evaluations across 11 benchmarks demonstrate the effectiveness and generalizability of SGL, achieving up to 91% pruning ratio for visual tokens while retaining competitive performance. The code is publicly available at https://github.com/NUS-HPC-AI-Lab/SGL. Wangbo Zhao, Yizeng Han, Jiasheng Tang, Zhikai Li, Yibing Song, Kai Wang 0036, Zhangyang Wang, Yang You 0001 |
CVPR | 1 |
| 2025 | EA-Vit: Efficient Adaptation for Elastic Vision TransformerabstractVision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming and energy-intensive. To address this issue, we propose an efficient ViT adaptation framework that enables a single adaptation process to generate multiple models of varying sizes for deployment on platforms with various resource constraints. Our approach comprises two stages. In the first stage, we enhance a pre-trained ViT with a nested elastic architecture that enables structural flexibility across MLP expansion ratio, number of attention heads, embedding dimension, and network depth. To preserve pre-trained knowledge and ensure stable adaptation, we adopt a curriculum-based training strategy that progressively increases elasticity. In the second stage, we design a lightweight router to select submodels according to computational budgets and downstream task demands. Initialized with Pareto-optimal configurations derived via a customized NSGA-II algorithm, the router is then jointly optimized with the backbone. Extensive experiments on multiple benchmarks demonstrate the effectiveness and versatility of EA-ViT. The code is available at https://github.com/zcxcf/EA-ViT. Wangbo Zhao, Yuhao Zhou 0004, Weidong Tang, Shuo Wang 0001, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Kai Wang 0036 |
ICCV | 2 |
| 2025 | Dynamic Diffusion TransformerabstractDiffusion Transformer (DiT), an emerging diffusion model for image generation,
has demonstrated superior performance but suffers from substantial computational
costs. Our investigations reveal that these costs stem from the static inference
paradigm, which inevitably introduces redundant computation in certain diffusion
timesteps and spatial regions. To address this inefficiency, we propose Dynamic
Diffusion Transformer (DyDiT), an architecture that dynamically adjusts its compu-
tation along both timestep and spatial dimensions during generation. Specifically,
we introduce a Timestep-wise Dynamic Width (TDW) approach that adapts model
width conditioned on the generation timesteps. In addition, we design a Spatial-
wise Dynamic Token (SDT) strategy to avoid redundant computation at unnecessary
spatial locations. Extensive experiments on various datasets and different-sized
models verify the superiority of DyDiT. Notably, with <3% additional fine-tuning it-
erations, our method reduces the FLOPs of DiT-XL by 51%, accelerates generation
by 1.73×, and achieves a competitive FID score of 2.07 on ImageNet. Wangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang 0036, Yibing Song, Gao Huang 0001, Fan Wang 0019, Yang You 0001 |
ICLR | 1 |
| 2025 | Unsupervised Learning for Class Distribution MismatchabstractClass distribution mismatch (CDM) refers to the discrepancy between class distributions in training data and target tasks. Previous methods address this by designing classifiers to categorize classes known during training, while grouping unknown or new classes into an "other" category. However, they focus on semi-supervised scenarios and heavily rely on labeled data, limiting their applicability and performance.
To address this, we propose Unsupervised Learning for Class Distribution Mismatch (UCDM), which constructs positive-negative pairs from unlabeled data for classifier training. Our approach randomly samples images and uses a diffusion model to add or erase semantic classes, synthesizing diverse training pairs. Additionally, we introduce a confidence-based labeling mechanism that iteratively assigns pseudo-labels to valuable real-world data and incorporates them into the training process.
Extensive experiments on three datasets demonstrate UCDM’s superiority over previous semi-supervised methods. Specifically, with a 60\% mismatch proportion on Tiny-ImageNet dataset, our approach, without relying on labeled data, surpasses OpenMatch (with 40 labels per class) by 35.1%, 63.7%, and 72.5% in classifying known, unknown, and new classes. Pan Du 0002, Wangbo Zhao, Xinai Lu, Zhikai Li, Chaoyu Gong, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Kai Wang 0036, Yang You 0001 |
ICML | 2 |
| 2025 | Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsabstractModern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of
prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to
\textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs.
We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research. Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036 |
NeurIPS | 6 |
| 2025 | Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning PathsabstractMamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields.
In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness.
Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34\% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5\% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages. Jindi Lv, Yuhao Zhou 0004, Mingjia Shi, Zhiyuan Liang, Xiaojiang Peng, Wangbo Zhao, Jiancheng Lv 0001, Kai Wang 0036 |
NeurIPS | 7 |
| 2025 | Scaling Up Parameter Generation: A Recurrent Diffusion ApproachabstractParameter generation has long struggled to match the scale of today's large vision and language models, curbing its broader utility. In this paper, we introduce Recurrent Diffusion for Large-Scale Parameter Generation (RPG), a novel framework that generates full neural network parameters—up to hundreds of millions—on a single GPU. Our approach first partitions a network's parameters into non-overlapping 'tokens', each corresponding to a distinct portion of the model. A recurrent mechanism then learns the inter-token relationships, producing 'prototypes' which serve as conditions for a diffusion process that ultimately synthesizes the full parameters. Across a spectrum of architectures and tasks—including ResNets, ConvNeXts and ViTs on ImageNet-1K and COCO, and even LoRA-based LLMs—RPG achieves performance on par with fully trained networks while avoiding excessive memory overhead. Notably, it generalizes beyond its training set to generate valid parameters for previously unseen tasks, highlighting its flexibility in dynamic and open-ended scenarios. By overcoming the longstanding memory and scalability barriers, RPG serves as a critical advance in 'AI generating AI', potentially enabling efficient weight generation at scales previously deemed infeasible. Kai Wang 0036, Dongwen Tang, Wangbo Zhao, Konstantin Schürholt, Zhangyang Wang, Yang You 0001 |
NeurIPS | 3 |
| 2025 | REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingabstractDiffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plateaus or even degrades performance later. We trace this failure to the capacity mismatch: once the generative student begins modeling the joint data distribution, the teacher's lower-dimensional embeddings and attention patterns become a straitjacket rather than a guide. We then introduce HASTE (Holistic Alignment with Stage-wise Termination for Efficient training), a two-phase schedule that keeps the help and drops the hindrance. Phase I applies a holistic alignment loss that simultaneously distills attention maps (relational priors) and feature projections (semantic anchors) from the teacher into mid-level layers of the DiT, yielding rapid convergence. Phase II then performs one-shot termination that deactivates the alignment loss, once a simple trigger such as a fixed iteration is hit, freeing the DiT to focus on denoising and exploit its generative capacity. HASTE speeds up training of diverse DiTs without architecture changes. On ImageNet 256×256, it reaches the vanilla SiT-XL/2 baseline FID in 50 epochs and matches REPA’s best FID in 500 epochs, amounting to a 28× reduction in optimization steps. HASTE also improves text-to-image DiTs on MS-COCO, proving to be a simple yet principled recipe for efficient diffusion training across various tasks. Wangbo Zhao, Yuhao Zhou 0004, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Kaipeng Zhang, Zhangyang Wang, Kai Wang 0036, Yang You 0001 |
NeurIPS | 2 |
| 2025 | Neural-Driven Image EditingabstractTraditional image editing typically relies on manual prompting, making it labor-intensive and inaccessible to individuals with limited motor control or language abilities. Leveraging recent advances in brain-computer interfaces (BCIs) and generative models, we propose LoongX, a hands-free image editing approach driven by multimodal neurophysiological signals.
LoongX utilizes state-of-the-art diffusion models trained on a comprehensive dataset of 23,928 image editing pairs, each paired with synchronized electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), photoplethysmography (PPG), and head motion signals that capture user intent.
To effectively address the heterogeneity of these signals, LoongX integrates two key modules. The cross-scale state space (CS3) module encodes informative modality-specific features. The dynamic gated fusion (DGF) module further aggregates these features into a unified latent space, which is then aligned with edit semantics via fine-tuning on a diffusion transformer (DiT).
Additionally, we pre-train the encoders using contrastive learning to align cognitive states with semantic intentions from embedded natural language.
Extensive experiments demonstrate that LoongX achieves performance comparable to text-driven methods (CLIP-I: 0.6605 vs. 0.6558; DINO: 0.4812 vs. 0.4637) and outperforms them when neural signals are combined with speech (CLIP-T: 0.2588 vs. 0.2549). These results highlight the promise of neural-driven generative models in enabling accessible, intuitive image editing and open new directions for cognitive-driven creative technologies. The code and dataset are released on the project website: https://loongx1.github.io. Xiaopeng Peng 0001, Wangbo Zhao, Zilong Ye, Suorong Yang, Jiadong Pan, Yuanxiang Chen, Kai Wang 0036, Xiaojun Chang, Gang Pan 0001, Shurong Dong, Kaipeng Zhang, Yang You 0001 |
NeurIPS | 4 |
| 2024 | VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt LearningabstractSalient object detection (SOD) and camouflaged object detection (COD) are related yet distinct binary mapping tasks. These tasks involve multiple modalities, sharing commonalities and unique cues. Existing research often employs intricate task-specific specialist models, potentially leading to redundancy and suboptimal results. We introduce VS-Code, a generalist model with novel 2D prompt learning, to jointly address four SOD tasks and three COD tasks. We utilize VST as the foundation model and introduce 2D prompts within the encoder-decoder architecture to learn domain and task-specific knowledge on two separate dimensions. A prompt discrimination loss helps disentangle peculiarities to benefit model optimization. VSCode outperforms state-of-the-art methods across six tasks on 26 datasets and exhibits zero-shot generalization to unseen tasks by combining 2D prompts, such as RGB-D COD. Source code has been available at https://github.com/Sssssuperior/VSCode. Nian Liu 0002, Wangbo Zhao, Xuguang Yang, Dingwen Zhang, Deng-Ping Fan, Fahad Shahbaz Khan, Junwei Han 0001 |
CVPR | 3 |
| 2024 | MMBench: Is Your Multi-modal Model an All-Around Player?
Yuan Liu 0025, Haodong Duan, Yuanhan Zhang, Bo Li 0080, Songyang Zhang 0001, Wangbo Zhao, Yike Yuan, Jiaqi Wang 0003, Conghui He, Ziwei Liu 0002, Kai Chen 0026, Dahua Lin |
ECCV (6) | 6 |
| 2024 | Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationabstractExisting parameter-efficient fine-tuning (PEFT) methods have achieved significant success on vision transformers (ViTs) adaptation by improving parameter efficiency. However, the exploration of enhancing inference efficiency during adaptation remains underexplored. This limits the broader application of pre-trained ViT models, especially when the model is computationally extensive. In this paper, we propose Dynamic Tuning (DyT), a novel approach to improve both parameter and inference efficiency for ViT adaptation. Specifically, besides using the lightweight adapter modules, we propose a token dispatcher to distinguish informative tokens from less important ones, allowing the latter to dynamically skip the original block, thereby reducing the redundant computation during inference. Additionally, we explore multiple design variants to find the best practice of DyT. Finally, inspired by the mixture-of-experts (MoE) mechanism, we introduce an enhanced adapter to further boost the adaptation performance. We validate DyT across various tasks, including image/video recognition and semantic segmentation. For instance, DyT achieves superior performance compared to existing PEFT methods while evoking only 71% of their FLOPs on the VTAB-1K benchmark. Wangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song, Kai Wang 0036, Gao Huang 0001, Fan Wang 0019, Yang You 0001 |
NeurIPS | 1 |
| 2024 | An Evidential Multi-Target Domain Adaptation Method Based on Weighted Fusion for Cross-Domain Pattern ClassificationabstractFor cross-domain pattern classification, the supervised information (i.e., labeled patterns) in the source domain is often employed to help classify the unlabeled target domain patterns. In practice, multiple target domains are usually available. The unlabeled patterns (in different target domains) which have high-confidence predictions, can also provide some pseudo-supervised information for the downstream classification task. The performance in each target domain would be further improved if the pseudo-supervised information in different target domains can be effectively used. To this end, we propose an evidential multi-target domain adaptation (EMDA) method to take full advantage of the useful information in the single-source and multiple target domains. In EMDA, we first align distributions of the source and target domains by reducing maximum mean discrepancy (MMD) and covariance difference across domains. After that, we use the classifier learned by the labeled source domain data to classify query patterns in the target domains. The query patterns with high-confidence predictions are then selected to train a new classifier for yielding an extra piece of soft classification results of query patterns. The two pieces of soft classification results are then combined by evidence theory. In practice, their reliabilities/weights are usually diverse, and an equal treatment of them often yields the unreliable combination result. Thus, we propose to use the distribution discrepancy across domains to estimate their weighting factors, and discount them before fusing. The evidential combination of the two pieces of discounted soft classification results is employed to make the final class decision. The effectiveness of EMDA was verified by comparing with many advanced domain adaptation methods on several cross-domain pattern classification benchmark datasets. Linqing Huang, Wangbo Zhao, Duo Yang 0008, Alan Wee-Chung Liew |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Learning Complementary Spatial-Temporal Transformer for Video Salient Object DetectionabstractBesides combining appearance and motion information, another crucial factor for video salient object detection (VSOD) is to mine spatial-temporal (ST) knowledge, including complementary long-short temporal cues and global-local spatial context from neighboring frames. However, the existing methods only explored part of them and ignored their complementarity. In this article, we propose a novel complementary ST transformer (CoSTFormer) for VSOD, which has a short-global branch and a long-local branch to aggregate complementary ST contexts. The former integrates the global context from the neighboring two frames using dense pairwise attention, while the latter is designed to fuse long-term temporal information from more consecutive frames with local attention windows. In this way, we decompose the ST context into a short-global part and a long-local part and leverage the powerful transformer to model the context relationship and learn their complementarity. To solve the contradiction between local window attention and object motion, we propose a novel flow-guided window attention (FGWA) mechanism to align the attention windows with object and camera movements. Furthermore, we deploy CoSTFormer on fused appearance and motion features, thus enabling the effective combination of all three VSOD factors. Besides, we present a pseudo video generation method to synthesize sufficient video clips from static images for training ST saliency models. Extensive experiments have verified the effectiveness of our method and illustrated that we achieve new state-of-the-art results on several benchmark datasets. Nian Liu 0002, Kepan Nan, Wangbo Zhao, Xiwen Yao, Junwei Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-grained Temporal Prototype Learning for Few-shot Video Object SegmentationabstractFew-Shot Video Object Segmentation (FSVOS) aims to segment objects in a query video with the same category defined by a few annotated support images. However, this task was seldom explored. In this work, based on IPMT, a state-of-the-art few-shot image segmentation method that combines external support guidance information with adaptive query guidance cues, we propose to leverage multi-grained temporal guidance information for handling the temporal correlation nature of video data. We decompose the query video information into a clip prototype and a memory prototype for capturing local and long-term internal temporal guidance, respectively. Frame prototypes are further used for each frame independently to handle fine-grained adaptive guidance and enable bidirectional clip-frame prototype communication. To reduce the influence of noisy memory, we propose to leverage the structural similarity relation among different predicted regions and the support for selecting reliable memory frames. Furthermore, a new segmentation loss is also proposed to enhance the category discriminability of the learned prototypes. Experimental results demonstrate that our proposed video IPMT model significantly outperforms previous models on two benchmark datasets. Code is available at https://github.com/nankepan/VIPMT. Nian Liu 0002, Kepan Nan, Wangbo Zhao, Yuanwei Liu, Xiwen Yao, Salman Khan 0001, Hisham Cholakkal, Rao Muhammad Anwer, Junwei Han 0001, Fahad Shahbaz Khan |
ICCV | 3 |
| 2023 | A new multi-source Transfer Learning method based on Two-stage Weighted Fusion
Linqing Huang, Jinfu Fan, Wangbo Zhao, Yang You 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationabstractText-based video segmentation aims to segment the target object in a video based on a describing sentence. Incorporating motion information from optical flow maps with appearance and linguistic modalities is crucial yet has been largely ignored by previous work. In this paper, we design a method to fuse and align appearance, motion, and linguistic features to achieve accurate segmentation. Specifically, we propose a multi-modal video transformer, which can fuse and aggregate multi-modal and temporal features between frames. Furthermore, we design a language-guided feature fusion module to progressively fuse appearance and motion features in each feature level with guidance from linguistic features. Finally, a multi-modal alignment loss is proposed to alleviate the semantic gap between features from different modalities. Extensive experiments on A2D Sentences and J-HMDB Sentences verify the performance and the generalization ability of our method compared to the state-of-the-art methods. Wangbo Zhao, Kai Wang 0036, Xiangxiang Chu, Fuzhao Xue, Xinchao Wang, Yang You 0001 |
CVPR | 1 |
| 2022 | Instance-Level Relative Saliency Ranking With Graph ReasoningabstractConventional salient object detection models cannot differentiate the importance of different salient objects. Recently, two works have been proposed to detect saliency ranking by assigning different degrees of saliency to different objects. However, one of these models cannot differentiate object instances and the other focuses more on sequential attention shift order inference. In this paper, we investigate a practical problem setting that requires simultaneously segment salient instances and infer their relative saliency rank order. We present a novel unified model as the first end-to-end solution, where an improved Mask R-CNN is first used to segment salient instances and a saliency ranking branch is then added to infer the relative saliency. For relative saliency ranking, we build a new graph reasoning module by combining four graphs to incorporate the instance interaction relation, local contrast, global contrast, and a high-level semantic prior, respectively. A novel loss function is also proposed to effectively train the saliency ranking branch. Besides, a new dataset and an evaluation metric are proposed for this task, aiming at pushing forward this field of research. Finally, experimental results demonstrate that our proposed model is more effective than previous methods. We also show an example of its practical usage on adaptive image retargeting. Nian Liu 0002, Long Li 0008, Wangbo Zhao, Junwei Han 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Weakly Supervised Video Salient Object DetectionabstractSignificant performance improvement has been achieved for fully-supervised video salient object detection with the pixel-wise labeled training datasets, which are time-consuming and expensive to obtain. To relieve the burden of data annotation, we present the first weakly super-vised video salient object detection model based on relabeled “fixation guided scribble annotations”. Specifically, an "Appearance-motion fusion module" and bidirectional ConvLSTM based framework are proposed to achieve effective multi-modal learning and long-term temporal context modeling based on our new weak annotations. Further, we design a novel foreground-background similarity loss to further explore the labeling similarity across frames. A weak annotation boosting strategy is also introduced to boost our model performance with a new pseudo-label generation technique. Extensive experimental results on six benchmark video saliency detection datasets illustrate the effectiveness of our solution1. Wangbo Zhao, Jing Zhang 0052, Long Li 0008, Nick Barnes, Nian Liu 0002, Junwei Han 0001 |
CVPR | 1 |
| 2021 | Light Field Saliency Detection with Dual Local Graph Learning and Reciprocative GuidanceabstractThe application of light field data in salient object detection is becoming increasingly popular recently. The difficulty lies in how to effectively fuse the features within the focal stack and how to cooperate them with the feature of the all-focus image. Previous methods usually fuse focal stack features via convolution or ConvLSTM, which are both less effective and ill-posed. In this paper, we model the information fusion within focal stack via graph networks. They introduce powerful context propagation from neighbouring nodes and also avoid ill-posed implementations. On the one hand, we construct local graph connections thus avoiding prohibitive computational costs of traditional graph networks. On the other hand, instead of processing the two kinds of data separately, we build a novel dual graph model to guide the focal stack fusion process using all-focus patterns. To handle the second difficulty, previous methods usually implement one-shot fusion for focal stack and all-focus features, hence lacking a thorough exploration of their supplements. We introduce a reciprocative guidance scheme and enable mutual guidance between these two kinds of information at multiple steps. As such, both kinds of features can be enhanced iteratively, finally benefiting the saliency prediction. Extensive experimental results show that the proposed models are all beneficial and we achieve significantly better results than state-of-the-art methods. Nian Liu 0002, Wangbo Zhao, Dingwen Zhang, Junwei Han 0001, Ling Shao 0001 |
ICCV | 2 |
| 2021 | SCG: Saliency and Contour Guided Salient Instance SegmentationabstractDifferent from conventional instance segmentation, salient instance segmentation (SIS) faces two difficulties. The first is that it involves segmenting salient instances only while ignoring background, and the second is that it targets generic object instances without pre-defined object categories. In this paper, based on the state-of-the-art Mask R-CNN model, we propose to leverage complementary saliency and contour information to handle these two challenges. We first improve Mask R-CNN by introducing an interleaved execution strategy and proposing a novel mask head network to incorporate global context within each RoI. Then we add two branches to Mask R-CNN for saliency and contour detection, respectively. We fuse the Mask R-CNN features with the saliency and contour features, where the former supply pixel-wise saliency information to help with identifying salient regions and the latter provide a generic object contour prior to help detect and segment generic objects. We also propose a novel multiscale global attention model to generate attentive global features from multiscale representative features for feature fusion. Experimental results demonstrate that all our proposed model components can improve SIS performance. Finally, our overall model outperforms state-of-the-art SIS methods and Mask R-CNN by more than 6% and 3%, respectively. By using additional multitask training data, we can further improve the model performance on the ILSO dataset. Nian Liu 0002, Wangbo Zhao, Ling Shao 0001, Junwei Han 0001 |
IEEE Trans. Image Process. | 2 |