Changwei Wang 0001

dblp:61/8501 · DBLP profile ↗
← Back
62ranked-venue papers
12as first author
62since 2021 · last 2026
0000-0001-8259-7717ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 8 first-author · 37 since 2021Artificial intelligence and machine learning · 29 · 6 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
abstract
Verifying the complex and multi-step reasoning of Large Language Models (LLMs) is a critical challenge, as holistic methods often overlook localized flaws. Step-by-step validation is a promising alternative, yet existing methods are often rigid. They struggle to adapt to diverse reasoning structures, from formal proofs to informal natural language narratives. To address this adaptability gap, we propose the Graph of Verification (GoV), a novel framework for adaptable and multi-granular verification. GoV's core innovation is its flexible node block architecture. This mechanism allows GoV to adaptively adjust its verification granularity—from atomic steps for formal tasks to entire paragraphs for natural language—to match the native structure of the reasoning process. This flexibility allows GoV to resolve the fundamental trade-off between verification precision and robustness. Experiments on both well-structured and loosely-structured benchmarks demonstrate GoV's versatility. The results show that GoV's adaptive approach significantly outperforms both holistic baselines and other state-of-the-art decomposition-based methods, establishing a new standard for training-free reasoning verification.
Jiwei Fang, Bin Zhang 0052, Changwei Wang 0001, Jin Wan, Zhiwei Xu 0005
AAAI3
2026 MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
abstract
Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality.
Run Ling, Ke Cao 0001, Ao Ma 0005, Runze He, Changwei Wang 0001, Rongtao Xu, Yihua Shao, Zhanjie Zhang, Guibing Guo, Jingjing Lv, Junjie Shen 0008, Ching Law, Xingwei Wang 0001
AAAI7
2026 WM-DETR: Dual-branch wavelet-Mamba and sparse attention for robust underwater object detection
Shunpeng Chen, Zenghuang Fu, Longzhao Huang, Shengpeng Xu, Yixian Kong, Changwei Wang 0001, Weiliang Meng, Xiaopeng Zhang 0001
Expert Syst. Appl.8
2026 Online knowledge distillation optimization based on Multi-Student model Multi-Task collaborative learning
Shibiao Xu, Shanshan Mo, Changwei Wang 0001, Hetong Wang, Rongtao Xu, Li Guo 0004
Knowl. Based Syst.4
2026 Tail-Aware Reconstruction of Incomplete Label Distributions With Low-Rank and Sparse Modeling
abstract
Label Distribution Learning (LDL) is a novel machine learning paradigm that addresses the problem of label ambiguity and has found widespread applications. However, obtaining complete label distributions in real-world scenarios is challenging, which has led to the emergence of Incomplete Label Distribution Learning (InLDL). Existing InLDL methods attempt to utilize low-rank label correlations to recover the complete label distribution. However, we find that real-world LDL datasets have animbalancednature; that is, the sum of the description degrees for normal labels is significantly larger than that for tail labels, which disrupts the low-rank assumption underlying the recovery of the label distribution. To solve the above problem, we propose Incomplete and Imbalance Label Distribution Learning (I2LDL), which makes the use of low-rank label correlations more reasonable for InLDL. Our method decomposes the recovered label distribution matrix into a low-rank component for frequent labels and a sparse component for tail labels, effectively capturing the structure of both head and tail labels. We further require that the entries in the observed positions of the recovered label distribution matrix be close to the observed values, and that the recovered label distribution for every instance forms a probability simplex (i.e., nonnegative entries summing to unity). Finally, the proposed model is optimized via the Alternating Direction Method of Multipliers (ADMM). We provide a theoretical analysis of its exact recovery guarantee under standard assumptions of incoherence, sparsity, and sufficient sampling. Furthermore, we establish a generalization error bound based on Rademacher complexity, offering theoretical insights into the learning performance of our method. Extensive experiments on 16 real-world datasets demonstrate the effectiveness and robustness of our framework compared to existing InLDL methods. The code is available at https://anonymous.4open.science/r/IncomLDL-tailaware-C021.
Zhiqiang Kou, Haoyuan Xuan, Hailin Wang 0001, Ming-Kun Xie, Changwei Wang 0001, Jing Wang 0113, Yuheng Jia, Xin Geng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Dark-EvGS: Event Camera as an Eye for Radiance Field in the Dark
abstract
In low-light environments, conventional cameras often struggle to capture clear multi-view images of objects due to dynamic range limitations and motion blur caused by long exposure. Event cameras, with their high-dynamic range and high-speed properties, have the potential to mitigate these issues. Additionally, 3D Gaussian Splatting (GS) enables radiance field reconstruction, facilitating bright frame synthesis from multiple viewpoints in low-light conditions. However, naively using an event-assisted 3D GS approach still faced challenges because, in low lights, events are noisy, frames lack quality, and the color tone may be inconsistent. To address these issues, we propose Dark-EvGS, the first event-assisted 3D GS framework that enables the reconstruction of bright frames from arbitrary viewpoints along the camera trajectory. Triplet-level supervision is proposed to gain holistic knowledge, granular details, and sharp scene rendering. The color tone matching block is proposed to guarantee the color consistency of the rendered frames. Furthermore, we introduce the first real-captured dataset for the event-guided bright frame synthesis task via 3D GS-based radiance field reconstruction. Experiments demonstrate that our method achieves better results than existing methods, conquering radiance field reconstruction under challenging low-light conditions. The code and sample data are included in the supplementary material.
Jingqian Wu, Peiqi Duan 0002, Zongqiang Wang, Changwei Wang 0001, Boxin Shi, Edmund Y. Lam
IEEE Trans. Image Process.4
2026 TeDri:Teacher-Driven Region Knowledge Distillation
abstract
Knowledge distillation as a practical tool to enhance the performance of small-capacity student networks on downstream tasks comes at the cost of a lengthy distillation process due to the online inference of teacher networks, especially when there is a large capacity gap between them. Therefore, in this paper, we propose a fast distillation framework called TeDri based on region images by offline saving relevant regional information and its teacher guidance. Specifically, first, to alleviate the lack of diversity caused by the fixed augmentation path in region images, we propose Teacher-driven MixUp strategies with mild intensity and advocate binding the mixing factor$\lambda$with teacher guidance confidence, where more confident category representations dominate the MixUp process. Furthermore, recognizing the need to evaluate these randomly cropped regions, and we propose region contrastive learning, encourage the student network to mimic the region partitioning behavior of the teacher, promoting a comprehensive understanding of global semantic content from multiple local perspectives. Finally, we introduce region mutual learning, employing spatial constraints among regions to require the student network towards consistent content interpretation across localized regions. Experiments on CIFAR-100 and ImageNet-1 K validate the effectiveness of the proposed TeDri, achieving competitive performance while significantly reducing training time.
Changwei Wang 0001, Rongtao Xu, Xingtian Pei, Shibiao Xu, Wenbo Xu 0003, Li Guo 0004
IEEE Trans. Knowl. Data Eng.2
2026 Adaptive in Adapter: Boosting Open-Vocabulary Semantic Segmentation With Adaptive Dropout Adapter
abstract
Open-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release theAdaptiveInAdapter (i.e.AIA) from the following two aspects:i)A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes.ii)A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available athttps://github.com/clearxu/AIA.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Jiguang Zhang, Xiaoqiang Teng, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.1
2026 Robust detection in complex construction sites: HiPA-DETR with weather-aware and cross-domain generalization
Zenghuang Fu, Muyang Zhang, Changwei Wang 0001, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Vis. Comput.6
2025 Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In addition, the lack of pixel-level correspondence supervision in the VPR dataset hinders further improvement of the local feature matching capability in the re-ranking stage. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency.
Changwei Wang 0001, Shunpeng Chen, Rongtao Xu, Jiguang Zhang, Haoran Yang 0003, Yu Zhang 0133, Kexue Fu 0001, Shide Du, Zhiwei Xu 0005, Longxiang Gao, Li Guo 0004, Shibiao Xu
AAAI1
2025 OpenViewer: Openness-Aware Multi-View Learning
abstract
Multi-view learning methods leverage multiple data sources to enhance perception by mining correlations across views, typically relying on predefined categories. However, deploying these models in real-world scenarios presents two primary openness challenges. 1) Lack of Interpretability: The integration mechanisms of multi-view data in existing black-box models remain poorly explained; 2) Insufficient Generalization: Most models are not adapted to multi-view scenarios involving unknown categories. To address these challenges, we propose OpenViewer, an openness-aware multi-view learning framework with theoretical support. This framework begins with a Pseudo-Unknown Sample Generation Mechanism to efficiently simulate open multi-view environments and previously adapt to potential unknown samples. Subsequently, we introduce an Expression-Enhanced Deep Unfolding Network to intuitively promote interpretability by systematically constructing functional prior-mapping modules and effectively providing a more transparent integration mechanism for multi-view data. Additionally, we establish a Perception-Augmented Open-Set Training Regime to significantly enhance generalization by precisely boosting confidences for known categories and carefully suppressing inappropriate confidences for unknown ones. Experimental results demonstrate that OpenViewer effectively addresses openness challenges while ensuring recognition performance for both known and unknown samples.
Shide Du, Zihan Fang 0002, Yanchao Tan, Changwei Wang 0001, Shiping Wang, Wenzhong Guo
AAAI4
2025 PanoDiT: Panoramic Videos Generation with Diffusion Transformer
abstract
As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-to-video (T2V) diffusion techniques have shown potential in standard video generation, they face challenges when applied to panoramic videos due to substantial differences in content and motion patterns. In this paper, we propose PanoDiT, a framework that utilizes the Diffusion Transformer (DiT) architecture to generate panoramic videos from text descriptions. Unlike traditional methods that rely on UNet-based denoising, our method leverages a transformer architecture for denoising, incorporating both temporal and global attention mechanisms. This ensures coherent frame generation and smooth motion transitions, offering distinct advantages in long-horizon generation tasks. To further enhance motion and consistency in the generated videos, we introduce DTM-LoRA and two panoramic-specific losses. Compared to previous methods, our PanoDiT achieves state-of-the-art performance across various evaluation metrics and user study, with code is available in the supplementary material.
Muyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang 0001, Weiliang Meng, Jianwei Guo 0003, Xiaopeng Zhang 0001
AAAI4
2025 Dual Focus-Attention Transformer for Robust Point Cloud Registration
abstract
Recently, coarse-to-fine methods for point cloud registration have achieved great success, but few works deeply explore the impact of feature interaction at both coarse and fine scales. By visualizing attention scores and correspondences, we find that existing methods fail to achieve effective feature aggregation at the two scales during the feature interaction. To tackle this issue, we propose a Dual Focus-Attention Transformer framework, which only focuses on points relevant to the current point for feature interaction, avoiding interactions with irrelevant points. For the coarse scale, we design a superpoint focus-attention transformer guided by sparse keypoints, which are selected from the neighborhood of superpoints. For the fine scale, we only perform feature interaction between the point sets that belong to the same superpoint. Experiments show that our method achieve the state-of-the-art performance on three standard benchmarks. The code and pre-trained models are available at https://github.com/fukexue/DFAT.git.
Kexue Fu 0001, Mingzhi Yuan, Changwei Wang 0001, Weiguang Pang, Jing Chi, Manning Wang, Longxiang Gao
CVPR3
2025 AASD: Accelerate Inference by Aligning Speculative Decoding in Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have achieved notable success in visual instruction tuning, yet their inference is time-consuming due to the auto-regressive decoding of Large Language Model (LLM) backbone. Traditional methods for accelerating inference, including model compression and migration from language model acceleration, often compromise output quality or face challenges in effectively integrating multimodal features. To address these issues, we propose AASD, a novel framework for Accelerating inference with refined KV Cache and Aligning speculative decoding in MLLMs. Our approach leverages the target model’s cached KeyValue (KV) pairs to extract vital information for generating draft tokens, enabling efficient speculative decoding. To reduce the computational burden associated with long multimodal token sequences, we introduce a KV Projector to compress the KV Cache while maintaining representational fidelity. Additionally, we design a Target-Draft Attention mechanism that optimizes the alignment between the draft model and the target model, achieving the benefits of real inference scenarios with minimal computational overhead. Extensive experiments on mainstream MLLMs demonstrate that our method achieves up to a $2 \times$ inference speedup without sacrificing accuracy. This study not only provides an effective and lightweight solution for accelerating MLLM inference but also introduces a novel alignment strategy for speculative decoding in multimodal contexts, laying a strong foundation for future research in efficient MLLMs. Code is availiable at https://github.com/transcend-0/ASD
Muyang Zhang, Weiguang Pang, Yuzhi Chen, Rongtao Xu, Kexue Fu 0001, Changwei Wang 0001, Longxiang Gao
DAC8
2025 CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation
abstract
Data-Free Knowledge Distillation (DFKD) enables the knowledge transfer from the given pre-trained teacher network to the target student model without access to the real training data. Existing DFKD methods focus primarily on improving image recognition performance on associated datasets, often neglecting the crucial aspect of the transferability of learned representations. In this paper, we propose Category-Aware Embedding Data-Free Knowledge Distillation (CAE-DFKD), which addresses at the embedding level the limitations of previous rely on image-level methods to improve model generalization but fail when directly applied to DFKD. The superiority and flexibility of CAE-DFKD are extensively evaluated, including: i.) Significant efficiency advantages resulting from altering the generator training paradigm; ii.) Competitive performance with existing DFKD state-of-the-art methods on image recognition tasks; iii.) Remarkable transferability of data-free learned representations demonstrated in downstream tasks.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Yu Zhang 0133, Jie Zhou 0001, Li Guo 0004
DAC2
2025 MVPS: Multi-View Adaptive Prompt Synergy for Zero-shot Anomaly Detection
abstract
Zero-shot anomaly detection (ZSAD) in industrial domains faces significant challenges due to the diverse manifestations of anomalies across scales and semantic levels. Existing methods, relying on single prompt spaces, struggle to generalize across these variations. So they exhibit limited generalization across scales and semantic levels. We propose Multi-View Adaptive Prompting Synergy (MVPS), a novel framework that establishes multiple scale-aware prompt spaces to enhance anomaly detection generalization. MVPS comprises three synergistic components: Hierarchical Multi-modal Prompt Tuning (HMPT) for generating scale-aware prompts, Dual-stream Prompt Tuning Orchestration (DPTO) for achieving robust cross-modal feature alignment, and Multi-View Prompt Composition Learning (MVPCL) for effective scale feature perception. This approach enables comprehensive capture anomaly feature across multiple semantic levels and scales, overcoming limitations of single-space representations and scale-insensitive feature alignment. Extensive experiments on seven benchmark datasets demonstrate that MVPS achieves state-of-the-art performance, exhibiting superior generalization capability across diverse anomaly categories and industrial domain. The code is available at https://github.com/MLY-0546/mvps.
Longzhao Huang, Changwei Wang 0001, Rongtao Xu, Shibiao Xu
ICME3
2025 Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
abstract
Camera-based occupancy prediction is a main-stream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and complex cascaded frameworks, with good yet limited performance. Few studies explore from the perspective of representation fusion, leaving the rich diversity of features in 2D images underutilized. Motivated by this, we propose CIGOcc, a two-stage occupancy prediction framework based on multi-level representation fusion. CIGOcc extracts segmentation, graphics, and depth features from an input image and introduces a deformable multi-level fusion mechanism to fuse these three multi-level features. Additionally, CIGOcc incorporates knowledge distilled from SAM to further enhance prediction accuracy. Without increasing training costs, CIGOcc achieves state-of-the-art performance on the SemanticKITTI benchmark. The code is provided in the supplementary material and will be released project page.
Rongtao Xu, Jinzhou Lin 0001, Jialei Zhou, Jiahua Dong 0001, Changwei Wang 0001, Ruisheng Wang 0001, Li Guo 0004, Shibiao Xu, Xiaodan Liang
ICRA5
2025 3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
abstract
With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by lever-aging the strengths of foundational models. The framework integrates key components, including multi-modal embedding, cross-modal interaction, and a language model decoder, to process natural language instructions and 3D scene data. This approach facilitates enhanced reasoning and response generation in complex 3D environments. Using the ScanNet 3D scene dataset, along with text annotations from ScanQA and ScanRefer, 3D-MoRe generates 62,000 question-answer (QA) pairs and 73,000 object descriptions across 1,513 scenes. We also employ various data augmentation techniques and implement semantic filtering to ensure high-quality data. Experiments on ScanQA demonstrate that 3D-MoRe significantly outperforms state-of-the-art baselines, with the CIDEr score improving by 2.15%. Similarly, on ScanRefer, our approach achieves a notable increase in [email protected] by 1.84%, highlighting its effectiveness in both tasks. Our code and generated datasets will be publicly released to benefit the community, and both can be accessed on the https://3D-MoRe.github.io.
Rongtao Xu, Mingming Yu, Dong An 0002, Shunpeng Chen, Changwei Wang 0001, Li Guo 0004, Xiaodan Liang, Shibiao Xu
IROS6
2025 LargeMvC-Net: Anchor-based Deep Unfolding Network for Large-scale Multi-view Clustering
abstract
Deep anchor-based multi-view clustering methods enhance the scalability of neural networks by utilizing representative anchors to reduce the computational complexity of large-scale clustering. Despite their scalability advantages, existing approaches often incorporate anchor structures in a heuristic or task-agnostic manner, either through post-hoc graph construction or as auxiliary components for message passing. Such designs overlook the core structural demands of anchor-based clustering, neglecting key optimization principles. To bridge this gap, we revisit the underlying optimization problem of large-scale anchor-based multi-view clustering and unfold its iterative solution into a novel deep network architecture, termed LargeMvC-Net. The proposed model decomposes the anchor-based clustering process into three modules: RepresentModule, NoiseModule, and AnchorModule, corresponding to representation learning, noise suppression, and anchor estimation. Each module is derived by unfolding a step of the original optimization procedure into a dedicated network component, providing structural clarity and optimization traceability. In addition, an unsupervised reconstruction loss aligns each view with the anchor-induced latent space, encouraging consistent clustering structures across views. Extensive experiments on several large-scale multi-view benchmarks show that LargeMvC-Net consistently outperforms state-of-the-art methods in terms of both effectiveness and scalability. The source data, code. https://github.com/dushide/LargeMvC-Net_ACMMM_2025, and extended version http://arxiv.org/abs/2507.20980 are available.
Shide Du, Zihan Fang 0002, Wendi Zhao, Yilin Wu 0001, Changwei Wang 0001, Shiping Wang
ACM Multimedia6
2025 Collaboration Wins More: Dual-Modal Collaborative Attention Reinforcement for Mitigating Large Vision Language Models Hallucination
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual-language understanding for downstream multimodal tasks. However, these models often generate descriptions containing objects or details not present in the input image, a phenomenon commonly referred to as ''hallucination''. Existing methods focus solely on single-side hallucination mitigation: Intra-modal-only reinforcement (e.g. visual attention enhancement) ignores prompt-based guidance; Inter-modal-only correlation correction may introduce low-information visual tokens to mislead reasoning. To tackle this challenge, we propose Dual-Modal Collaborative Attention Reinforcement (DuCAR). Specifically, DuCAR is equipped with intra-visual CLS-driven sampling and cross-modal dynamic sampling, extracting important visual tokens guided by intra- and inter-modal joint information. During the multimodal fusion stage, DuCAR adaptively enhances the attention weights of these visual tokens. Our sampling and enhancement strategies in DuCAR simultaneously reinforces informative visual tokens, and suppresses attention dispersion towards question-irrelevant visual information. We conduct extensive experiments on the POPE and CHAIR hallucination benchmarks, demonstrating that our method outperforms existing state-of-the-art mitigation baselines and effectively reduces hallucinations in text generated by LVLMs. The code is available in the https://github.com/xjy2020/DuCAR.
Jiye Xie, Liangliang You, Zhiqiang Kou, Kexue Fu 0001, Youyang Qu, Wenjie Yang 0005, Jianwei Guo 0003, Weiliang Meng, Longxiang Gao, Haoran Yang 0003, Changwei Wang 0001, Yu Zhang 0133
ACM Multimedia14
2025 Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning
abstract
Current text-to-image diffusion generation typically employs complete-text conditioning. Due to the intricate syntax, diffusion transformers (DiTs) inherently suffer from a comprehension defect of complete-text captions. One-fly complete-text input either overlooks critical semantic details or causes semantic confusion by simultaneously modeling diverse semantic primitive types. To mitigate this defect of DiTs, we propose a novel split-text conditioning framework named DiT-ST. This framework converts a complete-text caption into a split-text caption, a collection of simplified sentences, to explicitly express various semantic primitives and their interconnections. The split-text caption is then injected into different denoising stages of DiT-ST in a hierarchical and incremental manner. Specifically, DiT-ST leverages Large Language Models to parse captions, extracting diverse primitives and hierarchically sorting out and constructing these primitives into a split-text input. Moreover, we partition the diffusion denoising process according to its differential sensitivities to diverse semantic primitive types and determine the appropriate timesteps to incrementally inject tokens of diverse semantic primitive types into input tokens via cross-attention. In this way, DiT-ST enhances the representation learning of specific semantic primitive types across different stages. Extensive experiments validate the effectiveness of our proposed DiT-ST in mitigating the complete-text comprehension defect. Datasets and models are available.
Yu Zhang 0133, Jialei Zhou, Xinchen Li, Qi Zhang 0020, Zhongwei Wan, Duoqian Miao 0001, Changwei Wang 0001, Longbing Cao
NeurIPS7
2025 PointMM: A Hybrid Mamba-Transformer Framework for Point Cloud Analysis with Morton Reordering Strategy
Changwei Wang 0001, Shujun Gu, Chuanfu Wu, Longxiang Gao, Kexue Fu 0001, Youyang Qu
PRCV (10)2
2025 A Mamba-KAN Joint UNet Framework for Medical Image Segmentation
Haoyu Zhou, Changwei Wang 0001, Weiguang Pang, Lei Cui 0006, Shujun Gu, Longxiang Gao, Kexue Fu 0001, Youyang Qu
PRCV (3)2
2025 Dual prototypes contrastive learning based semi-supervised segmentation method for intelligent medical applications
Tianai Yue, Rongtao Xu, Jingqian Wu, Wenjie Yang 0005, Shide Du, Changwei Wang 0001
Eng. Appl. Artif. Intell.6
2025 FDBPL: Faster distillation-based prompt learning for region-aware vision-language models adaptation
Changwei Wang 0001, Rongtao Xu, Longzhao Huang, Wenbo Xu 0003, Li Guo 0004, Shibiao Xu
Expert Syst. Appl.3
2025 C2Fi-NeRF: Coarse to fine inversion NeRF for 6D pose estimation
Jiguang Zhang, Zhaohui Zhang 0002, Xuxiang Feng, Shibiao Xu, Rongtao Xu, Changwei Wang 0001, Kexue Fu 0001, Jiaxi Sun, Weilong Ding 0001
Expert Syst. Appl.6
2025 Disentangled Active Learning on Graphs
Haoran Yang 0003, Junli Wang 0001, Rui Duan 0003, Changwei Wang 0001, ChunGang Yan
Neural Networks4
2025 Generalization Boosted Adapter for Open-Vocabulary Segmentation
abstract
Vision-language models (VLMs) have demonstrated remarkable open-vocabulary object recognition capabilities, motivating their adaptation for dense prediction tasks like segmentation. However, directly applying VLMs to such tasks remains challenging due to their lack of pixel-level granularity and the limited data available for fine-tuning, leading to overfitting and poor generalization. To address these limitations, we propose Generalization Boosted Adapter (GBA), a novel adapter strategy that enhances the generalization and robustness of VLMs for open-vocabulary segmentation. GBA comprises two core components: (1) a Style Diversification Adapter (SDA) that decouples features into amplitude and phase components, operating solely on the amplitude to enrich the feature space representation while preserving semantic consistency; and (2) a Correlation Constraint Adapter (CCA) that employs cross-attention to establish tighter semantic associations between text categories and target regions, suppressing irrelevant low-frequency “noise” information and avoiding erroneous associations. Through the synergistic effect of the shallow SDA and the deep CCA, GBA effectively alleviates overfitting issues and enhances the semantic relevance of feature representations. As a simple, efficient, and plug-and-play component, GBA can be flexibly integrated into various CLIP-based methods, demonstrating broad applicability and achieving state-of-the-art performance on multiple open-vocabulary segmentation benchmarks. Code are available athttps://github.com/clearxu/BGA.
Changwei Wang 0001, Xuxiang Feng, Rongtao Xu, Longzhao Huang, Li Guo 0004, Shibiao Xu
IEEE Trans. Circuits Syst. Video Technol.2
2025 DFMC: Feature-Driven Data-Free Knowledge Distillation
abstract
Data-Free Knowledge Distillation (DFKD) enables knowledge transfer from teacher networks without access to the real dataset. However, generator-based DFKD methods often suffer from insufficient diversity or low-confidence in synthetic images, negatively impacting student network performance. This paper introduces DFMC, a generative feature-driven framework to mitigate the inherent limitations of DFKD. We propose exploiting semantic description between generative feature domains to guide augmentation strategies, avoiding random abstract inputs caused by inconsistent semantic quality. Then, by applying noise to the generative features, we produce contrastive learning pairs indirectly, limiting the sampling range of the feature domain to encourage the student network to learn domain-invariant features. Finally, we guide the student network to deeply mimic the teacher’s layer-wise implicit classification behavior for the augmented synthetic images. Extensive experiments across various datasets and downstream tasks demonstrate the effectiveness of DFMC, achieving significant improvements while preventing student networks from overfitting to semantic ambiguous images.
Rongtao Xu, Changwei Wang 0001, Shunpeng Chen, Shibiao Xu, Guangyuan Xu, Li Guo 0004
IEEE Trans. Circuits Syst. Video Technol.3
2025 Segment Anything Model Is a Good Teacher for Local Feature Learning
abstract
Local feature detection and description play an important role in many computer vision tasks, which are designed to detect and describe keypoints in any scene and any downstream task. Data-driven local feature learning methods need to rely on pixel-level correspondence for training. However, a vast number of existing approaches ignored the semantic information on which humans rely to describe image pixels. In addition, it is not feasible to enhance generic scene keypoints detection and description simply by using traditional common semantic segmentation models because they can only recognize a limited number of coarse-grained object classes. In this paper, we propose SAMFeat to introduce SAM (segment anything model), a foundation model trained on 11 million images, as a teacher to guide local feature learning. SAMFeat learns additional semantic information brought by SAM and thus is inspired by higher performance even with limited training samples. To do so, first, we construct an auxiliary task of Attention-weighted Semantic Relation Distillation (ASRD), which adaptively distillates feature relations with category-agnostic semantic information learned by the SAM encoder into a local feature learning network, to improve local feature description using semantic discrimination. Second, we develop a technique called Weakly Supervised Contrastive Learning Based on Semantic Grouping (WSC), which utilizes semantic groupings derived from SAM as weakly supervised signals, to optimize the metric space of local descriptors. Third, we design an Edge Attention Guidance (EAG) to further improve the accuracy of local feature detection and description by prompting the network to pay more attention to the edge region guided by SAM. SAMFeat's performance on various tasks, such as image matching on HPatches, and long-term visual localization on Aachen Day-Night showcases its superiority over previous local features. The release code is available at https://github.com/vignywang/SAMFeat.
Jingqian Wu, Rongtao Xu, Zach Wood-Doughty, Changwei Wang 0001, Shibiao Xu, Edmund Y. Lam
IEEE Trans. Image Process.4
2025 Token Masking Transformer for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is both a promising and challenging task that aims to achieve object localization exclusively through image category labels for supervision. Visual transformers have recently been applied to WSOL, demonstrating significant success through the exploitation of long-range feature dependencies in self-attention mechanisms. However, the transformer-based approach suffers from the same partial activation problem as the CNN-based approach due to the use of the classification task to train self-attention map, i.e., only a few discriminative regions are assigned high attention response and thus the localization map does not cover the whole object. To alleviate this problem, we propose a plug-and-play Token Masking Transformer (TMT) method to help transformer-based WSOL methods to obtain a more complete localization map by dynamic discriminative token masking. Specifically, a batch-wise discriminative token selection strategy is first introduced to flexibly determine the tokens to be masked in each image. Then, we design a token masking transformer block to perform token masking and inspire the network to mine more object-related tokens. Besides, we also design an intermediate token activation loss to further improve the performance of TMT by imposing constraints on intermediate tokens. Extensive experiments demonstrate that our TMT can substantially improve the performance of existing transformer-based methods without increasing the computational cost, and achieves state-of-the-art performance on two mainstream benchmarks.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Man Zhang 0005, Xiaopeng Zhang 0001
IEEE Trans. Multim.2
2025 OV-BIS: Open-Vocabulary Boundary Guide Zero-Shot 3D Instance Segmentation
abstract
Open vocabulary 3D instance segmentation aims to align 3D instance segmentation results with natural language text, thereby achieving semantic prediction without relying on predefined class labels for specific scenes, which has been widely used in the field of multimedia. Current open vocabulary 3D instance segmentation methods mainly rely on 2D masks provided by various 2D segmentation foundation models. However, in complex scenes, the calculation of 2D masks often struggles to balance over-segmentation of large objects and under-segmentation of small objects. In this paper, we introduce OV-BIS, a novel zero-shot open vocabulary 3D instance segmentation method that leverages instance boundary information to improve 3D semantic segmentation performance. The key insight of our method is that the edge map as 3D boundary projection is suitable for multi-scale tasks and capable of compensating for the weakness of 2D masks in multi-scale adaptability for complex scenes. Our method aggregates multiview edge maps and 2D masks, iteratively guiding the merging of over-segmented point clouds with regions growing to cluster 3D primitives into distinct 3D instances. By projecting 3D instances onto images and using CLIP to calculate semantic features from multiple perspectives with an outliers filter, 3D semantic instance segmentation has been achieved. Experiments on multiple datasets demonstrate the superiority of our method.
Tinghao Yi, Shaohu Wang, Zhengtao Zhang, Changwei Wang 0001, Dong-Ming Yan 0001, Rongtao Xu, Enhong Chen
IEEE Trans. Multim.4
2025 SRIF: Data-Free Knowledge Distillation via Stable Regulation and Input Filtering
abstract
Data-free knowledge distillation (DFKD) enables knowledge transfer from a pre-trained teacher to a student network without accessing the real dataset. However, generator-based DFKD methods struggle to ensure that the synthetic images accurately reflect the real dataset distribution. The update of the generator network relies heavily on teacher category guidance, but varying teacher prediction accuracy across categories leads to inconsistent synthetic image quality. Such variations introduce a distribution shift between synthetic and real datasets, negatively impacting student network performance during knowledge distillation. To address this challenge, we propose the SRIF, comprising two components: Student-Driven Flexible Filtering (SDFF) and Re-weighting for Independent Regularization (RIR). SDFF filters out synthetic images affected by the category distribution shift during data generation, producing a more reliable dataset. RIR, applied during distillation, encourages the student to learn stable causal relationships through sample reweighting. Both components flexibly integrate into existing DFKD frameworks, improving performance while reducing training costs.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Jie Zhou 0001, Longxiang Gao, Wenbo Xu 0003, Li Guo 0004
IEEE Trans. Multim.3
2025 PDFT: parameter-diminish fine-tuning for transformer-based models
Muyang Zhang, Weiliang Meng, Mingda Jia, Jiaming Gu, Yihua Shao, Changwei Wang 0001, Rongtao Xu, Xiaopeng Zhang 0001
Vis. Comput.6
2024 Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
abstract
Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and incorporate Visual Prompt Training (VPT) to uphold CLIP's generalization capacity, they still fall short in fully harnessing CLIP's potential for pixel-level unseen class demarcation and precise pixel predictions. To further stimulate CLIP's zero-shot dense prediction capability, we propose SPT-SEG, a one-stage approach that improves CLIP's adaptability from image to pixel. Specifically, we initially introduce Spectral Prompt Tuning (SPT), incorporating spectral prompts into the CLIP visual encoder's shallow layers to capture structural intricacies of images, thereby enhancing comprehension of unseen classes. Subsequently, we introduce the Spectral Guided Decoder (SGD), utilizing both high and low-frequency information to steer the network's spatial focus towards more prominent classification features, enabling precise pixel-level prediction outcomes. Through extensive experiments on two public datasets, we demonstrate the superiority of our method over state-of-the-art approaches, performing well across all classes and particularly excelling in handling unseen classes.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Li Guo 0004, Man Zhang 0005, Xiaopeng Zhang 0001
AAAI3
2024 AC-CAM: Affinity-Aware Contrast CAM for Weakly-Supervised Semantic Segmentation on MRI Brain Tumor
abstract
Applying the latest visual transformer (ViT) to Weakly-Supervised Semantic Segmentation (WSSS) can compensate for the local perception limitations of CNN, but it also brings about the over-smoothing problem, that is, the final patch labels tend to be uniform. To overcome this challenge, we present an Affinity-Aware Contrast Class Activation Maps (AC-CAM) framework aimed at enhancing WSSS for MRI Brain Tumor analysis by exploiting only image-level labels. We propose two main components: the Affinity-Aware Token Contrast Module (ATCM) and the Affinity-Aware Refine Module (ARM). ATCM utilizes semantic affinities from attention maps to improve the contrast between patch tokens, effectively reducing the over-smoothing tendency of Vision Transformers (ViT). ARM refines the pseudo labels further, incorporating RGB and affinity information to capture the intricate details of the target objects. Our approach capitalizes on the global feature capturing capabilities of ViT, producing more accurate pseudo-labels for WSSS. The framework is optimized through a composite loss function that ensures the consistency of representations for positive token pairs and discriminability for negative ones. Experiments show that our method achieves state-of-the-art performance on the BraTS 2021 dataset.
Jingqian Wu, Changwei Wang 0001, Duzhen Zhang, Rongtao Xu
BIBM3
2024 MIM-HD: Making Smaller Masked Autoencoder Better with Efficient Distillation
abstract
Self-supervised learning and knowledge distillation intersect to achieve exceptional performance on downstream tasks across diverse network capacities. This paper introduces MIM-HD, which implements enhancements for masked image modeling (MIM) distillation, in two key aspects. First, a vision transformer head-level relation adaptive distillation approach is proposed, allowing the student to dynamically draw multi-source knowledge from the teacher based on its evolving state, compatible with scenarios where teacher-student transformer block head count differs. Second, to address the overemphasis on the encoder and neglect of the decoder role in maintaining representation consistency in previous MIM distillations, a dual-view decoding strategy for latent visual representations is introduced, reusing the teacher’s decoder to alleviate MIM burdens on smaller networks. MIM-HD effectiveness is demonstrated through evaluations on ADE20K (mIoU) and ImageNet-1K (Acc), achieving +1.4% and +0.5% improved performance, respectively, compared to state-of-the-art methods, with substantial advantages on smaller pre-training datasets. Moreover, MIM-HD achieves superior efficiency, reducing pre-training epochs from 300 to 100.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Li Guo 0004, Jiguang Zhang, Xiaoqiang Teng, Wenbo Xu 0003
ECAI2
2024 HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
abstract
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few pixels. However, it encounters difficulties due to the diminutive size of the objects and the generally complex backgrounds in infrared images. In this paper, we propose a deep learning method, HCF-Net, that significantly improves infrared small object detection performance through multiple practical modules. Specifically, it includes the parallelized patch-aware attention (PPA) module, dimension-aware selective integration (DASI) module, and multi-dilated channel refiner (MDCR) module. The PPA module uses a multi-branch feature extraction strategy to capture feature information at different scales and levels. The DASI module enables adaptive channel selection and fusion. The MDCR module captures spatial features of different receptive field ranges through multiple depth-separable convolutional layers. Extensive experimental results on the SIRST infrared single-frame image dataset show that the proposed HCF-Net performs well, surpassing other traditional and deep learning models. Code is available at https://github.com/zhengshuchen/HCFNet.
Shibiao Xu, ShuChen Zheng, Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Xiaoqiang Teng, Ao Li 0002, Li Guo 0004
ICME5
2024 MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization
abstract
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, leading to rapid advancements in multimodal studies. However, CLIP faces a notable challenge in terms of *inefficient data utilization*. It relies on a single contrastive supervision for each image-text pair during representation learning, disregarding a substantial amount of valuable information that could offer richer supervision. Additionally, the retention of non-informative tokens leads to increased computational demands and time costs, particularly in CLIP's ViT image encoder. To address these issues, we propose **M**ulti-Perspective **L**anguage-**I**mage **P**retraining (**MLIP**). In MLIP, we leverage the frequency transform's sensitivity to both high and low-frequency variations, which complements the spatial domain's sensitivity limited to low-frequency variations only. By incorporating frequency transforms and token-level alignment, we expand CILP's single supervision into multi-domain and multi-level supervision, enabling a more thorough exploration of informative image features. Additionally, we introduce a token merging method guided by comprehensive semantics from the frequency and spatial domains. This allows us to merge tokens to multi-granularity tokens with a controllable compression rate to accelerate CLIP. Extensive experiments validate the effectiveness of our design.
Yu Zhang 0133, Qi Zhang 0020, Zixuan Gong, Yiwei Shi, Duoqian Miao 0001, Kun Yi 0001, Wei Fan 0010, Liang Hu 0004, Changwei Wang 0001
ICML12
2024 DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic Segmentation
abstract
The complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored and face two key inherent difficulties: over-attention and noise from different modal data. To overcome these challenges, we propose a Deformable Multimodal Representation Fusion (DefFusion) framework consisting mainly of a Deformable Representation Fusion Transformer and Dynamic Representation Augmentation Modules. The Deformable Representation Fusion Transformer introduces the deformable mechanism in multimodal fusion, avoiding over-attention and improving efficiency by adaptively modeling a 2D key/value set for a given 3D query, thus enabling multimodal fusion with higher flexibility. To enhance the 2D representation and 3D representation, the Dynamic Representation Enhancement Module is proposed to dynamically remove noise in the input representation via Dynamic Grouped Representation Generation and Dynamic Mask Generation. Extensive experiments validate that our model achieves the best 3D semantic segmentation performance on SemanticKITTI and NuScenes benchmarks.
Rongtao Xu, Changwei Wang 0001, Duzhen Zhang, Man Zhang 0005, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICRA2
2024 NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
abstract
Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips.
Zixuan Gong, Guangyin Bao, Qi Zhang 0020, Zhongwei Wan, Duoqian Miao 0001, Shoujin Wang, Lei Zhu 0003, Changwei Wang 0001, Rongtao Xu, Liang Hu 0004, Yu Zhang 0133
NeurIPS8
2024 Learning adaptive shift and task decoupling for discriminative one-step person search
Qixian Zhang, Duoqian Miao 0001, Qi Zhang 0020, Changwei Wang 0001, Hongyun Zhang 0001, Cairong Zhao
Knowl. Based Syst.4
2024 DomainFeat: Learning Local Features With Domain Adaptation
abstract
Accurate and efficient keypoint detection and description is a fundamental step in various computer vision tasks. In this paper, we extract robust descriptors and detect accurate keypoints by learning local Features with Domain adaptation (DomainFeat). Specifically, our Domainfeat includes image-level domain invariance supervision, pixel-level domain consistency supervision, Pixel-Adaptive keypoint Detection(PA-Det), and cross-domain dataset with domain stable point supervision. First, we introduce the image-level domain invariance supervision to make the high-level feature distributions from different domains close by fusing domain-invariant representations in the decoder. Furthermore, to compensate for the inconsistency between descriptors corresponding to the keypoints at the pixel level, we propose the pixel-level domain consistency supervision. Then we present the Pixel-Adaptive keypoint Detection to efficiently detect accurate keypoints, which can improve accuracy by enhancing the local consistency of heatmaps. Finally, we propose an efficient approach to construct data and supervision labels in diverse domains, which can tackle complex application scenarios. With these novel modules and supervision methods, our DomainFeat can make feature detectors more accurate and descriptors more robust. Extensive experiments confirm that Domainfeat achieves state-of-the-art performance on benchmarks such as Aachen-Day-Night localization, HPatches image matching, and the challenging DNIM dataset.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Exploring Intrinsic Discrimination and Consistency for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is a challenging and promising task that aims to localize objects solely based on the supervision of image category labels. In the absence of annotated bounding boxes, WSOL methods must employ the intrinsic properties of the image classification task pipeline to generate object localizations. In this work, we propose a WSOL method for exploring the Intrinsic Discrimination and Consistency in the image classification task pipeline, and call it as IDC. First, we develop a Triplet Metrics Based Foreground Modeling (TMFM) framework to directly predict object foreground regions using intrinsic discrimination. Unlike Class Activation Map (CAM) based methods that also rely on intrinsic discrimination, our TMFM framework alleviates the problem of only focusing on the most discriminative parts by optimizing foreground and background regions synergistically. Second, we design a Dual Geometric Transformation Consistency Constraints (DGTC2) training strategy to introduce additional supervision and regularization constraints for WSOL by leveraging intrinsic geometric transformation consistency. The proposed pixel-wise and object-wise consistency constraint losses cost-effectively provide spontaneous supervision for WSOL. Extensive experiments show that our IDC method achieves significant and consistent performance gains compared to existing state-of-the-art WSOL approaches. Code is available at: https://github.com/vignywang/IDC.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Xiaopeng Zhang 0001
IEEE Trans. Image Process.1
2024 SkinFormer: Learning Statistical Texture Representation With Transformer for Skin Lesion Segmentation
abstract
Accurate skin lesion segmentation from dermoscopic images is of great importance for skin cancer diagnosis. However, automatic segmentation of melanoma remains a challenging task because it is difficult to incorporate useful texture representations into the learning process. Texture representations are not only related to the local structural information learned by CNN, but also include the global statistical texture information of the input image. In this paper, we propose a transFormer network (SkinFormer) that efficiently extracts and fuses statistical texture representation for Skin lesion segmentation. Specifically, to quantify the statistical texture of input features, a Kurtosis-guided Statistical Counting Operator is designed. We propose Statistical Texture Fusion Transformer and Statistical Texture Enhance Transformer with the help of Kurtosis-guided Statistical Counting Operator by utilizing the transformer's global attention mechanism. The former fuses structural texture information and statistical texture information, and the latter enhances the statistical texture of multi-scale features. Extensive experiments on three publicly available skin lesion datasets validate that our SkinFormer outperforms other SOAT methods, and our method achieves 93.2% Dice score on ISIC 2018. It can be easy to extend SkinFormer to segment 3D images in the future.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE J. Biomed. Health Informatics2
2024 PSTNet: Enhanced Polyp Segmentation With Multi-Scale Alignment and Frequency Domain Integration
abstract
Accurate segmentation of colorectal polyps in colonoscopy images is crucial for effective diagnosis and management of colorectal cancer (CRC). However, current deep learning-based methods primarily rely on fusing RGB information across multiple scales, leading to limitations in accurately identifying polyps due to restricted RGB domain information and challenges in feature misalignment during multi-scale aggregation. To address these limitations, we propose the Polyp Segmentation Network with Shunted Transformer (PSTNet), a novel approach that integrates both RGB and frequency domain cues present in the images. PSTNet comprises three key modules: the Frequency Characterization Attention Module (FCAM) for extracting frequency cues and capturing polyp characteristics, the Feature Supplementary Alignment Module (FSAM) for aligning semantic information and reducing misalignment noise, and the Cross Perception localization Module (CPM) for synergizing frequency cues with high-level semantics to achieve efficient polyp segmentation. Extensive experiments on challenging datasets demonstrate PSTNet's significant improvement in polyp segmentation accuracy across various metrics, consistently outperforming state-of-the-art methods. The integration of frequency domain cues and the novel architectural design of PSTNet contribute to advancing computer-assisted polyp segmentation, facilitating more accurate diagnosis and management of CRC.
Rongtao Xu, Changwei Wang 0001, Xiuli Li, Shibiao Xu, Li Guo 0004
IEEE J. Biomed. Health Informatics3
2024 Wave-Like Class Activation Map With Representation Fusion for Weakly-Supervised Semantic Segmentation
abstract
The Class Activation Map (CAM) is widely used to generate pseudo-labels for Weakly Supervised Semantic Segmentation (WSSS), while it does not adequately consider the modeling of foreground-independent information, resulting in prone to false positive pixels. In this paper, we propose a Wave-like Class Activation Map (WaveCAM) from the perspective of representation fusion and dynamic aggregation representation to alleviate the above problem. Specifically, our WaveCAM includes the foreground-aware representation modeling that enhances perception of foreground information, and the foreground-independent representation modeling that enhances perception of foreground-independent information, and a representation-adaptive fusion module that fuses the two representations. Both representations are expressed as wave functions with amplitude and phase to dynamically aggregate representations and extract semantic information after initialization, and they are fused through the adaptive fusion module to obtain an output containing rich semantic information. Extensive experiments on PASCAL VOC 2012 dataset and MS COCO 2014 dataset validate that our WaveCAM can easily embed multi-stage WSSS and end-to-end WSSS, achieving the state-of-the-art performance.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.2
2024 Accurate Lung Nodule Segmentation With Detailed Representation Transfer and Soft Mask Supervision
abstract
Accurate lung lesion segmentation from computed tomography (CT) images is crucial to the analysis and diagnosis of lung diseases, such as COVID-19 and lung cancer. However, the smallness and variety of lung nodules and the lack of high-quality labeling make the accurate lung nodule segmentation difficult. To address these issues, we first introduce a novel segmentation mask named " soft mask," which has richer and more accurate edge details description and better visualization, and develop a universal automatic soft mask annotation pipeline to deal with different datasets correspondingly. Then, a novel network with detailed representation transfer and soft mask supervision (DSNet) is proposed to process the input low-resolution images of lung nodules into high-quality segmentation results. Our DSNet contains a special detailed representation transfer module (DRTM) for reconstructing the detailed representation to alleviate the small size of lung nodules images and an adversarial training framework with soft mask for further improving the accuracy of segmentation. Extensive experiments validate that our DSNet outperforms other state-of-the-art methods for accurate lung nodule segmentation, and has strong generalization ability in other accurate medical segmentation tasks with competitive results. Besides, we provide a new challenging lung nodules segmentation dataset for further studies (https://drive.google.com/file/d/15NNkvDTb_0Ku0IoPsNMHezJRTH1Oi1wm/view?usp=sharing).
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic Segmentation
abstract
Efficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high training efficiency. However, current methods suffer from insufficient extraction of comprehensive semantic information, resulting in low-quality pseudo-labels and sub-optimal solutions for end-to-end WSSS. To this end, we propose a simple and novel Self Correspondence Distillation (SCD) method to refine pseudo-labels without introducing external supervision. Our SCD enables the network to utilize feature correspondence derived from itself as a distillation target, which can enhance the network's feature learning process by complementing semantic information. In addition, to further improve the segmentation accuracy, we design a Variation-aware Refine Module to enhance the local consistency of pseudo-labels by computing pixel-level variation. Finally, we present an efficient end-to-end Transformer-based framework (TSCD) via SCD and Variation-aware Refine Module for the accurate WSSS task. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate that our method significantly outperforms other state-of-the-art methods. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/SCD-AAAI2023.
Rongtao Xu, Changwei Wang 0001, Jiaxi Sun, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
AAAI2
2023 Treating Pseudo-labels Generation as Image Matting for Weakly Supervised Semantic Segmentation
abstract
Generating accurate pseudo-labels under the supervision of image categories is a crucial step in Weakly Supervised Semantic Segmentation (WSSS). In this work, we propose a Mat-Label pipeline that provides a fresh way to treat WSSS pseudo-labels generation as an image matting task. By taking a trimap as input which specifies the foreground, background and unknown regions, the image matting task outputs an object mask with fine edges. The intuition behind our Mat-Label is that generating trimap is much easier than generating pseudo-labels directly under weakly supervised setting. Although current CAM-based methods are off-the-shelf solutions for generating a trimap, they suffer from cross-category and foreground-background pixel prediction confusion. To solve this problem, we develop a Double Decoupled Class Activation Map (D2CAM) for Mat-Label to generate a high-quality trimap. By drawing on the idea of metric learning, we explicitly model class activation map with category decoupling and foreground-background decoupling. We also design two simple yet effective refinement constraints for D2CAM to stabilize optimization and eliminate non-exclusive activation. Extensive experiments validate that our Mat-Label achieves substantial and consistent performance gains compared to current state-of-the-art WSSS approaches.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICCV1
2023 Automatic polyp segmentation via image-level and surrounding-level context fusion deep neural network
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.1
2023 Dual-stream Representation Fusion Learning for accurate medical image segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.2
2023 Attention Weighted Local Descriptors
abstract
Local features detection and description are widely used in many vision applications with high industrial and commercial demands. With large-scale applications, these tasks raise high expectations for both the accuracy and speed of local features. Most existing studies on local features learning focus on the local descriptions of individual keypoints, which neglect their relationships established from global spatial awareness. In this paper, we present AWDesc with a consistent attention mechanism (CoAM) that opens up the possibility for local descriptors to embrace image-level spatial awareness in both the training and matching stages. For local features detection, we adopt local features detection with feature pyramid to obtain more stable and accurate keypoints localization. For local features description, we provide two versions of AWDesc to cope with different accuracy and speed requirements. On the one hand, we introduce Context Augmentation to address the inherent locality of convolutional neural networks by injecting non-local context information, so that local descriptors can "look wider to describe better". Specifically, well-designed Adaptive Global Context Augmented Module (AGCA) and Diverse Surrounding Context Augmented Module (DSCA) are proposed to construct robust local descriptors with context information from global to surrounding. On the other hand, we design an extremely lightweight backbone network coupled with the proposed special knowledge distillation strategy to achieve the best trade-off in accuracy and speed. What is more, we perform thorough experiments on image matching, homography estimation, visual localization, and 3D reconstruction tasks, and the results demonstrate that our method surpasses the current state-of-the-art local descriptors. Code is available at: https://github.com/vignywang/AWDesc.
Changwei Wang 0001, Rongtao Xu, Ke Lu 0002, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road Shapes
abstract
Automatically extracting roads from very high resolution (VHR) remote sensing images is of great importance in a wide range of remote sensing applications. However, complex shapes of roads (i.e., long, geometrically deformed, and thin) always affected the extraction accuracy, which is one of the challenges of road extraction. Based on the insight into road shape characteristics, we propose a novel road shape aware network (RSANet) to achieve efficient and accurate road extraction. First, we introduce the Efficient Strip Transformer Module (ESTM) to efficiently capture the global context to model the long-distance dependence required by the long roads. Second, we design a Geometric Deformation Estimation Module (GDEM) to adaptively extract the context from the shape deformation caused by shooting roads from different perspectives. Third, we provide a simple but effective Road Edge Focal Loss (REF loss) to make the network focus on optimizing the pixels around the road to alleviate the unbalanced distribution of foreground and background pixels caused by the roads being too thin. Finally, we conduct extensive evaluations on public datasets to verify the effectiveness of RSANet and each of the proposed components. Experiments validate that our RSANet outperforms state-of-the-art methods for road extraction in remote sensing images.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Jiguang Zhang, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 RSSFormer: Foreground Saliency Enhancement for Remote Sensing Land-Cover Segmentation
abstract
High spatial resolution (HSR) remote sensing images contain complex foreground-background relationships, which makes the remote sensing land cover segmentation a special semantic segmentation task. The main challenges come from the large-scale variation, complex background samples and imbalanced foreground-background distribution. These issues make recent context modeling methods sub-optimal due to the lack of foreground saliency modeling. To handle these problems, we propose a Remote Sensing Segmentation framework (RSSFormer), including Adaptive TransFormer Fusion Module, Detail-aware Attention Layer and Foreground Saliency Guided Loss. Specifically, from the perspective of relation-based foreground saliency modeling, our Adaptive Transformer Fusion Module can adaptively suppress background noise and enhance object saliency when fusing multi-scale features. Then our Detail-aware Attention Layer extracts the detail and foreground-related information via the interplay of spatial attention and channel attention, which further enhances the foreground saliency. From the perspective of optimization-based foreground saliency modeling, our Foreground Saliency Guided Loss can guide the network to focus on hard samples with low foreground saliency responses to achieve balanced optimization. Experimental results on LoveDA datasets, Vaihingen datasets, Potsdam datasets and iSAID datasets validate that our method outperforms existing general semantic segmentation methods and remote sensing segmentation methods, and achieves a good compromise between computational overhead and accuracy. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/RSSFormer-TIP2023.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Image Process.2
2023 CNDesc: Cross Normalization for Local Descriptors Learning
abstract
For a long time, the local descriptors learning benefited from the use of L2 normalization, which projects the descriptor space onto the hypersphere. However, there is no free lunch in the world. Although hypersphere description space stabilizes the optimization and improves the repeatability of the descriptors, it causes the descriptors to have a denser distribution, which reduces the discrimination between descriptors and leads to some incorrect matches. To alleviate this problem, we propose the learnablecross normalizationtechnology as an alternative to L2 normalization, which can achieve a consistent improvement in several of the current popular local descriptors. In addition, we propose an ER-Backbone that can efficiently reuse features in descriptors extraction and an IDC Loss that can provide an image-level description space distribution consistency constraint to further stimulate the performance of the local descriptors. Based on the above innovations, we provide a novel local descriptors extraction method named CNDesc. We perform experiments on image matching, homography estimation, 3D reconstruction, and visual localization tasks, and the results demonstrate that our CNDesc surpasses the current state-of-the-art local descriptors. Our code is available athttps://github.com/vignywang/CNDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.1
2022 MTLDesc: Looking Wider to Describe Better
abstract
Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe better'' by learning local Descriptors with More Than Local information (MTLDesc). Specifically, we resort to context augmentation and spatial attention mechanism to make the descriptors obtain non-local awareness. First, Adaptive Global Context Augmented Module and Diverse Local Context Augmented Module are proposed to construct robust local descriptors with context information from global to local. Second, we propose the Consistent Attention Weighted Triplet Loss to leverage spatial attention awareness in both optimization and matching of local descriptors. Third, Local Features Detection with Feature Pyramid is proposed to obtain more stable and accurate keypoints localization. With the above innovations, the performance of the proposed MTLDesc significantly surpasses the current state-of-the-art local descriptors on HPatches, Aachen Day-Night localization and InLoc indoor localization benchmarks. Our code is available at https://github.com/vignywang/MTLDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
AAAI1
2022 DOMAINDESC: Learning Local Descriptors With Domain Adaptation
abstract
Robust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustness of our DomainDesc by punishing inconsistent high-level feature distributions of different images, while we present the pixel-level cross-domain consistency loss to compensate for the inconsistency between the descriptors corresponding to the keypoints at the pixel level. Besides, we adopt a new architecture to make the descriptor contain as much information as possible, and combine triplet loss and cross-domain consistency loss for descriptor supervision to ensure the distinguished ability of our descriptor. Finally, we give a cross-domain dataset generation strategy to quickly construct our training dataset for diverse domains to adapt to complex application scenarios. Experiments validate that our DomainDesc achieves state-of-the-art performances on HPatches image matching benchmark and Aachen-Day-Night localization benchmark.
Rongtao Xu, Changwei Wang 0001, Bin Fan 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICASSP2
2022 Softgan: Towards Accurate Lung Nodule Segmentation via Soft Mask Supervision
abstract
Accurate lung nodule segmentation from Computed Tomog-raphy (CT) images is crucial to the analysis and diagnosis of lung diseases such as COVID-19 and lung cancer. How-ever, due to the variety of lung nodules and the lack of high-quality labeling, accurate lung nodule segmentation is still a challenging problem. In this paper, we propose a novel paradigm including an automatic accurate annotation pipeline and a segmentation network for this task. First, we introduce a new segmentation mask representation named Soft Mask which has richer and more accurate edge details description and better visualization, and we design a universal automatic Soft Mask annotation pipeline to deal with different datasets. Besides, we provide a new challenging lung nodules segmen-tation dataset with traditional binarized masks and our soft masks for further studies. Second, we propose an effective network called SoftGAN that includes an improved back-bone and an adversarial training framework with Soft Mask, in order to improve the performance of accurate lung nodules segmentation. Extensive experiments validate that our Soft-GAN outperforms the state-of-the-art methods for accurate lung nodule segmentation. [Datasetrelease]
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Qimin Peng, Xiaopeng Zhang 0001
ICME1
2022 DA-Net: Dual Branch Transformer and Adaptive Strip Upsampling for Retinal Vessels Segmentation
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (2)1
2022 Instance segmentation of biological images using graph convolutional network
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.3
2021 DC-Net: Dual Context Network for 2D Medical Image Segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (1)2