Yijiang Li

dblp:115/8640 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
26since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 5 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Towards Adversarially Robust Dataset Distillation by Curvature Regularization
abstract
Dataset distillation (DD) allows datasets to be distilled to fractions of their original size while preserving the rich distributional information so that models trained on the distilled datasets can achieve a comparable accuracy while saving significant computational loads. Recent research in this area has been focusing on improving the accuracy of models trained on distilled datasets. In this paper, we aim to explore a new perspective of DD. We study how to embed adversarial robustness in distilled datasets, so that models trained on these datasets maintain the high accuracy and meanwhile acquire better adversarial robustness. We propose a new method that achieves this goal by incorporating curvature regularization into the distillation process with much less computational overhead than standard adversarial training. Extensive empirical experiments suggest that our method not only outperforms standard adversarial training on both accuracy and robustness with less computation overhead but is also capable of generating robust distilled datasets that can withstand various adversarial attacks.
Eric Xue 0002, Yijiang Li, Haoyang Liu 0001, Peiran Wang, Haohan Wang
AAAI2
2025 Evaluating Vision Language Models Through Concept Hacking
Yijiang Li, Bingyang Wang, Tianwei Zhao, Qingying Gao, Hokin Deng, Dezhi Luo
CogSci1
2025 Probing Mechanical Reasoning in Large Vision Language Models
Yijiang Li, Qingying Gao, Haiyun Lyu, Dezhi Luo, Hokin Deng
CogSci2
2025 Probing Perceptual Constancy in Large Vision Language Models
Suyang Yu, Yijiang Li, Qingying Gao, Haiyun Lyu, Hokin Deng, Dezhi Luo
CogSci3
2025 Robust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-On
abstract
This paper tackles the emerging challenge of multi-view virtual try-on, utilizing both front- and back-view clothing images as inputs. Extending frontal try-on methods to a multi-view context is not straightforward. Simply concatenating the two input views or encoding their features for a generative model, such as a diffusion model, often fails to produce satisfactory results. The main challenge lies in effectively extracting and fusing meaningful clothing features from these input views. Existing explicit warping-based methods, which establish direct correspondence between input and target views, tend to introduce artifacts, particularly when there is a significant disparity between the input and target views. Conversely, implicit encoding-based methods often lose spatial information about clothing, resulting in outputs that lack detail. To overcome these challenges, we propose Robust-MVTON, an end-to-end method for robust and high-quality multi-view try-ons. Our approach introduces a novel cross-pose feature alignment technique to guide the fusion of clothing features and incorporates a newly designed loss function for training. With the fused multi-scale clothing features, we employ a coarse-to-fine diffusion model to generate realistic and detailed results. Extensive experiments conducted on the Deepfashion and MPV datasets affirm the superiority of our method, achieving state-of-the-art performance.
Yijiang Li, Dong Du 0002, Zheng Chong, Zhengwentai Sun, Jianhao Zeng, Yusheng Dai, Zhengyu Xie, Hairui Zhu, Xiaoguang Han 0001
CVPR2
2025 VideoOrion: Tokenizing Object Dynamics in Videos
abstract
We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos - the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to extract object dynamics through a detect-segment-track pipeline, encoding them into a set of object tokens by aggregating spatial-temporal object features. Our method addresses the persistent challenge in Video-LLMs of efficiently compressing high-dimensional video data into semantic tokens that are comprehensible to LLMs. Compared to prior methods which resort to downsampling the original video or aggregating visual tokens using resamplers, leading to information loss and entangled semantics, VideoOrion not only offers a more natural and efficient way to derive compact, disentangled semantic representations but also enables explicit object modeling of video content with minimal computational cost. Moreover, the introduced object tokens naturally allow VideoOrion to accomplish video-based referring tasks. Experimental results show that VideoOrion can learn to make good use of the object tokens, and achieves competitive results on both general video question answering and video-based referring benchmarks.
Yicheng Feng, Yijiang Li, Wanpeng Zhang 0002, Sipeng Zheng, Hao Luo 0011, Zihao Yue, Zongqing Lu 0002
ICCV2
2025 Dataset Distillation via the Wasserstein Metric
abstract
Dataset Distillation (DD) aims to generate a compact synthetic dataset that enables models to achieve performance comparable to training on the full large dataset, significantly reducing computational costs. Drawing from optimal transport theory, we introduce WMDD (Wasserstein Metric-based Dataset Distillation), a straightforward yet powerful method that employs the Wasserstein metric to enhance distribution matching. We compute the Wasserstein barycenter of features from a pretrained classifier to capture essential characteristics of the original data distribution. By optimizing synthetic data to align with this barycenter in feature space and leveraging per-class BatchNorm statistics to preserve intra-class variations, WMDD maintains the efficiency of distribution matching approaches while achieving state-of-the-art results across various high-resolution datasets. Our extensive experiments demonstrate WMDD's effectiveness and adaptability, highlighting its potential for advancing machine learning applications at scale.
Haoyang Liu 0001, Yijiang Li, Tiancheng Xing, Peiran Wang, Vibhu Dalal, Luwei Li, Jingrui He, Haohan Wang
ICCV2
2025 Unified Multimodal Understanding via Byte-Pair Visual Encoding
abstract
Multimodal large language models (MLLMs) have made significant progress in vision-language understanding, yet effectively aligning different modalities remains a fundamental challenge. We present a framework that unifies multimodal understanding by applying byte-pair encoding to visual tokens. Unlike conventional approaches that rely on modality-specific encoders, our method directly incorporates structural information into visual tokens, mirroring successful tokenization strategies in text-only language models. We introduce a priority-guided encoding scheme that considers both frequency and spatial consistency, coupled with a multi-stage training procedure based on curriculum-driven data composition. These enhancements enable the transformer model to better capture cross-modal relationships and reason with visual information. Comprehensive experiments demonstrate improved performance across diverse vision-language tasks. By bridging the gap between visual and textual representations, our approach contributes to the advancement of more capable and efficient multimodal foundation models.
Wanpeng Zhang 0002, Yicheng Feng, Hao Luo 0011, Yijiang Li, Zihao Yue, Sipeng Zheng, Zongqing Lu 0002
ICCV4
2025 From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
abstract
Multimodal Large Language Models have made significant strides in integrating visual and textual information, yet they often struggle with effectively aligning these modalities. We introduce a novel image tokenizer that bridges this gap by applying the principle of Byte-Pair Encoding (BPE) to visual data. Unlike conventional approaches that rely on separate visual encoders, our method directly incorporates structural prior information into image tokens, mirroring the successful tokenization strategies used in text-only Large Language Models. This innovative approach enables Transformer models to more effectively learn and reason across modalities. Through theoretical analysis and extensive experiments, we demonstrate that our BPE Image Tokenizer significantly enhances MLLMs' multimodal understanding capabilities, even with limited training data. Leveraging this method, we develop Being-VL-0, a model that demonstrates superior performance across various benchmarks and shows promising scalability, potentially paving the way for more efficient and capable multimodal foundation models. For further details, visit our website https://github.com/BeingBeyond/Being-VL-0.
Wanpeng Zhang 0002, Zilong Xie, Yicheng Feng, Yijiang Li, Xingrun Xing, Sipeng Zheng, Zongqing Lu 0002
ICLR4
2025 Core Knowledge Deficits in Multi-Modal Language Models
abstract
While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem from the absence of core knowledge—rudimentary cognitive abilities innate to humans from early childhood. To explore the core knowledge representation in MLLMs, we introduce CoreCognition, a large-scale benchmark encompassing 12 core knowledge concepts grounded in developmental cognitive science. We evaluate 230 models with 11 different prompts, leading to a total of 2,530 data points for analysis. Our experiments uncover four key findings, collectively demonstrating core knowledge deficits in MLLMs: they consistently underperform and show reduced, or even absent, scalability on low-level abilities relative to high-level ones. Finally, we propose Concept Hacking, a novel controlled evaluation method, that reveals MLLMs fail to progress toward genuine core knowledge understanding, but instead rely on shortcut learning as they scale. Project page at https://williamium3000.github.io/core-knowledge/.
Yijiang Li, Qingying Gao, Tianwei Zhao, Bingyang Wang, Haiyun Lyu, Robert D. Hawkins, Nuno Vasconcelos, Tal Golan, Dezhi Luo, Hokin Deng
ICML1
2025 EgoPrivacy: What Your First-Person Camera Says About You?
abstract
While the rapid proliferation of wearable cameras has raised significant concerns about egocentric video privacy, prior work has largely overlooked the unique privacy threats posed to the camera wearer. This work investigates the core question: How much privacy information about the camera wearer can be inferred from their first-person view videos? We introduce EgoPrivacy, the first large-scale benchmark for the comprehensive evaluation of privacy risks in egocentric vision. EgoPrivacy covers three types of privacy (demographic, individual, and situational), defining seven tasks that aim to recover private information ranging from fine-grained (e.g., wearer's identity) to coarse-grained (e.g., age group). To further emphasize the privacy threats inherent to egocentric vision, we propose Retrieval-Augmented Attack, a novel attack strategy that leverages ego-to-exo retrieval from an external pool of exocentric videos to boost the effectiveness of demographic privacy attacks. An extensive comparison of the different attacks possible under all threat models is presented, showing that private information of the wearer is highly susceptible to leakage. For instance, our findings indicate that foundation models can effectively compromise wearer privacy even in zero-shot settings by recovering attributes such as identity, scene, gender, and race with 70–80% accuracy. Our code and data are available at https://github.com/williamium3000/ego-privacy.
Yijiang Li, Genpei Zhang, Yi Li 0051, Xiaojun Shan, Dashan Gao 0001, Jiancheng Lyu, Ning Bi, Nuno Vasconcelos
ICML1
2025 FedSpaLLM: Federated Pruning of Large Language Models
abstract
Guangji Bai, Yijiang Li, Zilinghan Li, Liang Zhao, Kibaek Kim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Guangji Bai, Yijiang Li, Zilinghan Li, Liang Zhao 0002, Kibaek Kim
NAACL (Long Papers)2
2025 FedCleanse: Cleanse the backdoor attacks in federated learning system
Siquan Huang, Yijiang Li, Chong Chen 0011, Leyu Shi, Wentian Cai, Ying Gao 0004
Knowl. Based Syst.2
2025 FedID: Enhancing Federated Learning Security Through Dynamic Identification
abstract
Federated learning (FL), recognized for its decentralized and privacy-preserving nature, faces vulnerabilities to backdoor attacks that aim to manipulate the model's behavior on attacker-chosen inputs. Most existing defenses based on statistical differences take effect only against specific attacks. This limitation becomes significantly pronounced when malicious gradients closely resemble benign ones or the data exhibits non-IID characteristics, making the defenses ineffective against stealthy attacks. This paper revisits distance-based defense methods and uncovers two critical insights: First, Euclidean distance becomes meaningless in high dimensions. Second, a single metric cannot identify malicious gradients with diverse characteristics. As a remedy, we propose FedID, a simple yet effective strategy employing multiple metrics with dynamic weighting for adaptive backdoor detection. Besides, we present a modified z-score approach to select the gradients for aggregation. Notably, FedID does not rely on predefined assumptions about attack settings or data distributions and minimally impacts benign performance. We conduct extensive experiments on various datasets and attack scenarios to assess its effectiveness. FedID consistently outperforms previous defenses, particularly excelling in challenging Edge-case PGD scenarios. Our experiments highlight its robustness against adaptive attacks tailored to break the proposed defense and adaptability to a wide range of non-IID data distributions without compromising benign performance.
Siquan Huang, Yijiang Li, Chong Chen 0011, Ying Gao 0004, Xiping Hu
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Scope: On Detecting Constrained Backdoor Attacks in Federated Learning
abstract
Federated learning (FL) allows multiple clients to train an efficient deep-learning model collaboratively but is susceptible to backdoor attacks. Traditional detection-based defenses depend on specific metrics to distinguish client gradients. Defense-aware attackers exploit this by constraining attack gradients on these metrics to evade detection, leading to metric-constrained attacks. This paper concretely instantiates such threats and introduces cosine-constrained attacks, which successfully compromise advanced defenses based on cosine distance. To address the aforementioned challenge, we propose Scope, a novel defense that detects cosine-constrained attacks using cosine distance by exposing the constrained backdoor dimensions of attack gradients. Scope employs dimension-wise normalization and differential scaling to amplify the distinction between backdoor dimensions and benign or unused ones, countering sophisticated attackers’ attempts to obscure them. Moreover, we develop a novel clustering approach, namely Dominant Gradient Clustering (DGC), to isolate and eliminate backdoor gradients. Extensive experiments across various datasets, models, FL settings, and adversary scenarios demonstrate that Scope consistently outperforms existing defenses by a significant margin, especially against the cosine-constrained attack. Additionally, we present a Scope-tailored attack designed to evade Scope, but it remains ineffective even when maximizing stealthiness, further underscoring the robustness of Scope. We release our source code at:https://github.com/siquanhuang/Scope.
Siquan Huang, Yijiang Li, Xingfu Yan, Ying Gao 0004, Chong Chen 0011, Leyu Shi, Wing W. Y. Ng
IEEE Trans. Inf. Forensics Secur.2
2025 Enhancing Weakly Supervised Semantic Segmentation With Multi-Label Contrastive Learning and LLM Features Guidance
abstract
Histopathological whole-slide images (WSIs) segmentation is essential for precise tissue characterization in medical diagnostics. However, traditional approaches require labor-intensive pixel-level annotations. To this end, we study weakly supervised semantic segmentation (WSSS) which uses patch-level classification labels, reducing annotation efforts significantly. However, the complexity of WSIs and the challenge of sparse classification labels hinder effective dense pixel predictions. Moreover, due to the multi-label nature of WSI, existing approaches of single-label contrastive learning designed for the representation of single-category, neglecting the presence of other relevant categories and thus fail to adapt to WSI tasks. This paper presents a novel multi-label contrastive learning method for WSSS by incorporating class-specific embedding extraction with LLM features guidance. Specifically, we propose to obtain class-specific embeddings by utilizing classifier weights, followed by a dot-product-based attention fusion method that leverages LLM features to enrich their semantics, facilitating contrastive learning between different classes from single image. Besides, we propose a Robust Learning approach that leverages multi-layer features to evaluate the uncertainty of pseudo-labels, thereby mitigating the impact of noisy pseudo-labels on the learning process of segmentation. Extensive experiments have been conducted on two histopathological image segmentation datasets, i.e. LUAD dataset and BCSS dataset, demonstrating the effectiveness of our methods with leading performance.
Wentian Cai, Yijiang Li, Yandan Chen, G. Thippa Reddy, Wei Wang 0077, Ying Gao 0004
IEEE J. Biomed. Health Informatics2
2024 Improving Prompt-based News Recommendation with Individual Template and Customized Answer
abstract
Prompt learning plays a key role in aligning the task of news recommendation (NR) with the Pre-trained Language Models (PLMs). However, current prompt-based NR methods utilize fixed templates and answer words, ignoring the personalization of user's demand and the diversity between news topics. To this end, we propose an Automatic Prompt based NR (AutoPNR) scheme, which automatically generates individual templates for users according to their potential interests, and customized answer words w.r.t. the topics of candidate news. Concretely, such an individual template utilizes several specific tokens to encode a user's interest extracted from her/his reading history, while a pair of customized answer words are retrieved from a large vocabulary (often existing alongside PLMs) based on the topic of candidate news. Through extensive experiments on the real-world datasets, we show that our AutoPNR works well with different PLMs, and considerably outperforms state-of-the-art NR techniques.
Yijiang Li, Jun Wu 0007
CIKM1
2024 SparseLLM: Towards Global Pruning of Pre-trained Language Models
abstract
The transformative impact of large language models (LLMs) like LLaMA and GPT on natural language processing is countered by their prohibitive computational demands. Pruning has emerged as a pivotal compression strategy, introducing sparsity to enhance both memory and computational efficiency. Yet, traditional global pruning is impractical for LLMs due to scalability issues, while local pruning, despite its efficiency, leads to suboptimal solutions. Addressing these challenges, we propose *SparseLLM*, a novel framework that redefines the global pruning process into manageable, coordinated subproblems, allowing for resource-efficient optimization with global optimality. SparseLLM's approach, which conceptualizes LLMs as a chain of modular functions and leverages auxiliary variables for problem decomposition, not only facilitates a pragmatic application on LLMs but also demonstrates significant performance improvements, particularly in high-sparsity regimes where it surpasses current state-of-the-art methods. Our source code is publicly available at https://github.com/BaiTheBest/SparseLLM.
Guangji Bai, Yijiang Li, Chen Ling 0003, Kibaek Kim, Liang Zhao 0002
NeurIPS2
2024 A reformulation-enumeration MINLP algorithm for gas network design
Yijiang Li, Santanu Subhas Dey, Nikolaos V. Sahinidis
J. Glob. Optim.1
2024 Stacked Graph Fusion Denoising Autoencoder for Hyperspectral Anomaly Detection
abstract
Anomaly detection for hyperspectral images (HSIs) is a challenging problem to distinguish a few anomalous pixels from a majority of background pixels. Most existing methods cannot simultaneously explore both structural and spatial information from global and local perspectives. In this letter, we propose a stacked graph fusion denoising autoencoder (SGFDAE) for hyperspectral anomaly detection. Specifically, the global and local graphs are constructed from an HSI to explore potential structural and spatial information. With the designed graph fusion strategy, an advanced graph denoising autoencoder with deep architecture is developed in a hierarchical manner. To achieve better reconstruction and detection, a greedy layerwise unsupervised pretraining strategy is presented for network training. Experiments show that SGFDAE achieves 97.17%, 98.43%, and 98.90% detection accuracies by averaging the results of the datasets from three different scenes and outperforms the state-of-the-art methods.
Yongshan Zhang, Yijiang Li, Xinxin Wang 0003, Xinwei Jiang, Yicong Zhou
IEEE Geosci. Remote. Sens. Lett.2
2023 Consistent-Teacher: Towards Reducing Inconsistent Pseudo-Targets in Semi-Supervised Object Detection
abstract
In this study, we dive deep into the inconsistency of pseudo targets in semi-supervised object detection (SSOD). Our core observation is that the oscillating pseudo-targets undermine the training of an accurate detector. It injects noise into the student's training, leading to severe overfitting problems. Therefore, we propose a systematic solution, termed Consistent-Teacher, to reduce the inconsistency. First, adaptive anchor assignment (ASA) substitutes the static IoU-based strategy, which enables the student network to be resistant to noisy pseudo-bounding boxes. Then we calibrate the subtask predictions by designing a 3D feature alignment module (FAM-3D). It allows each classification feature to adaptively query the optimal feature vector for the regression task at arbitrary scales and locations. Lastly, a Gaussian Mixture Model (GMM) dynamically revises the score threshold of pseudo-bboxes, which stabilizes the number of ground truths at an early stage and remedies the unreliable supervision signal during training. Consistent-Teacher provides strong results on a large range of SSOD evaluations. It achieves 40.0 mAP with ResNet-50 backbone given only 10% of annotated MS-COCO data, which surpasses previous base-lines using pseudo labels by around 3 mAP. When trained on fully annotated MS-COCO with additional unlabeled data, the performance further increases to 47.7 mAP. Our code is available at https://github.com/Adamdad/ConsistentTeacher.
Xinjiang Wang, Xingyi Yang, Yijiang Li, Litong Feng, Shijie Fang, Chengqi Lyu, Kai Chen 0002, Wayne Zhang 0001
CVPR4
2023 Multi-metrics adaptively identifies backdoors in Federated learning
abstract
The decentralized and privacy-preserving nature of federated learning (FL) makes it vulnerable to backdoor attacks aiming to manipulate the behavior of the resulting model on specific adversary-chosen inputs. However, most existing defenses based on statistical differences take effect only against specific attacks, especially when the malicious gradients are similar to benign ones or the data are highly non-independent and identically distributed (non-IID). In this paper, we revisit the distance-based defense methods and discover that i) Euclidean distance becomes meaningless in high dimensions and ii) malicious gradients with diverse characteristics cannot be identified by a single metric. To this end, we present a simple yet effective defense strategy with multi-metrics and dynamic weighting to identify backdoors adaptively. Furthermore, our novel defense has no reliance on predefined assumptions over attack settings or data distributions and little impact on benign performance. To evaluate the effectiveness of our approach, we conduct comprehensive experiments on different datasets under various attack settings, where our method achieves the best defensive performance. For instance, we achieve the lowest backdoor accuracy of 3.06% under the most difficult Edge-case PGD, showing significant superiority over previous defenses. The experiments also demonstrate that our method can be well-adapted to a wide range of non-IID degrees without sacrificing the benign performance.
Siquan Huang, Yijiang Li, Chong Chen 0011, Leyu Shi, Ying Gao 0004
ICCV2
2023 Diverse Cotraining Makes Strong Semi-Supervised Segmentor
abstract
Deep co-training has been introduced to semi-supervised segmentation and achieves impressive results, yet few studies have explored the working mechanism behind it. In this work, we revisit the core assumption that supports co-training: multiple compatible and conditionally independent views. By theoretically deriving the generalization upper bound, we prove the prediction similarity between two models negatively impacts the model’s generalization ability. However, most current co-training models are tightly coupled together and violate this assumption. Such coupling leads to the homogenization of networks and confirmation bias which consequently limits the performance. To this end, we explore different dimensions of co-training and systematically increase the diversity from the aspects of input domains, different augmentations and model architectures to counteract homogenization. Our Diverse Co-training outperforms the state-of-the-art (SOTA) methods by a large margin across different evaluation protocols on the Pascal and Cityscapes. For example, we achieve the best mIoU of 76.2%, 77.7% and 80.2% on Pascal with only 92, 183 and 366 labeled images, surpassing the previous best results by more than 5%.
Yijiang Li, Xinjiang Wang, Lihe Yang, Litong Feng, Wayne Zhang 0001, Ying Gao 0004
ICCV1
2023 Improved YOLOv7 Based on Transformer for Object Detection in UAV-Captured Images
abstract
As the drone captures image targets at different flying altitudes, their scales may vary significantly, which can pose challenges for the object detection model to accurately detect them. Additionally, tiny objects in the image contain minimal information, making them difficult to distinguish from the background. To overcome these two challenges, we proposed a network architecture that aims to improve the accuracy of tiny object detection in drone images. Specially, we designed a tiny object detector(TOD) that can effectively extract features of tiny objects and distinguish between tiny object features and image background. Furthermore, this TOD module contains a Convolutional Visual Attention Network (CVAN) to better focus on the regions of tiny objects. Experimental results demonstrate that the proposed method achieves [email protected] accuracy of 53.9% on the VisDrone2021-test-dev dataset and improves by 2.8 % compared to YOLOv7.
Yuefan Luo, Qing Zhu 0003, Zhen Zhou 0003, Lin Chen 0034, Tianjian Jiang, Yijiang Li, Danwei Wang, Yaonan Wang 0001
SMC7
2023 DFTNet: Dual-Path Feature Transfer Network for Weakly Supervised Medical Image Segmentation
abstract
Medical image segmentation has long suffered from the problem of expensive labels. Acquiring pixel-level annotations is time-consuming, labor-intensive, and relies on extensive expert knowledge. Bounding box annotations, in contrast, are relatively easy to acquire. Thus, in this paper, we explore to segment images through a novel Dual-path Feature Transfer design with only bounding box annotations. Specifically, a Target-aware Reconstructor is proposed to extract target-related features by reconstructing the pixels within the bounding box through the channel and spatial attention module. Then, a sliding Feature Fusion and Transfer Module (FFTM) fuses the extracted features from Reconstructor and transfers them to guide the Segmentor for segmentation. Finally, we present the Confidence Ranking Loss (CRLoss) which dynamically assigns weights to the loss of each pixel based on the network's confidence. CRLoss mitigates the impact of inaccurate pseudo-labels on performance. Extensive experiments demonstrate that our proposed model achieves state-of-the-art performance on the Medical Segmentation Decathlon (MSD) Brain Tumour and PROMISE12 datasets.
Wentian Cai, Linsen Xie, Weixian Yang, Yijiang Li, Ying Gao 0004, Tingting Wang 0006
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 More than Encoder: Introducing Transformer Decoder to Upsample
abstract
Medical image segmentation methods downsample images for feature extraction and then upsample them to restore resolution for pixel-level predictions. In such schema, upsample technique is vital in restoring information for better performance. However, existing upsample techniques leverage little information from downsampling paths. The local and detailed feature from the shallower layer such as boundary and tissue texture is crucial in segmentation, especially medical image segmentation. To this end, we propose a novel upsample approach for medical image segmentation, Window Attention Upsample (WAU), which upsamples features conditioned on local and detailed features from downsampling path in local windows by introducing attention decoders of Transformer. WAU could serve as a general upsample method and be incorporated into any segmentation model that possesses lateral connections. We first propose the Attention Upsample which consists of Attention Decoder (AD) and bilinear upsample. AD leverages pixel-level attention to model longrange dependency and global information for a better upsample. Bilinear upsample is introduced as the residual connection to complement the upsampled features. Moreover, considering the extensive memory and computation cost of pixel-level attention, we further design a window attention scheme to restrict attention computation in local windows instead of the global range. We evaluate our method (WAU) on classic UNet structure with lateral connections and achieve state-of-the-art performance on Medical Segmentation Decathlon (MSD) Brain and Automatic Cardiac Diagnosis Challenge (ACDC) datasets. We also validate the effectiveness of our method on multiple classic architectures and achieve consistent improvement.
Yijiang Li, Wentian Cai, Ying Gao 0004, Chengming Li 0004, Xiping Hu
BIBM1