EDBT 2026 Demo / reviewers in the wild / expert
Zhijian Song
dblp:23/3491
· DBLP profile ↗
43ranked-venue papers
0as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image Content Matters: An Image Content Aware State Space Model for Accelerated MRI ReconstructionabstractThe challenge of accelerated MRI reconstruction lies in recovering high-quality images from undersampled k-space. Recently, the selective state space model (Mamba) has shown promising results in various tasks with balanced global receptive field and computational efficiency, shedding new light on MRI reconstruction. However, existing approaches directly flatten 2D images based on spatial positions and apply Mamba to vision tasks, failing to preserve and explore the content properties. In this paper, we posit that the key to unlocking Mamba's full potential for MRI reconstruction lies in content-aware sequence modeling. We investigate two fundamental challenges: (1) how to reasonably preserve semantic information when converting 2D images into 1D sequences, and (2) how to effectively identify and recover the crucial high-frequency textures. To this end, we introduce CAM, a novel framework that shifts Mamba-based MRI reconstruction from position-based to content-aware sequence modeling. Specifically, we introduce three modules: (1) the Semantic Preservation Scanning Module (SPSM) introduces learnable clustering centers to group similar pixels, establishing the semantic preserved sequence. (2) The Texture Extraction Scanning Module (TESM) acts as a differentiable local texture descriptor to estimate crucial high-frequency information, forming the texture emphasized sequence. (3) The Texture Enhancement Mamba Module (TEMM) further modulates the semantic sequence with texture-informed system matrices derived from the texture sequence, yielding both context- and texture-aware sequential representations. With these enhancements, CAM significantly outperforms existing methods across various datasets and under-sampling masks. Yucong Meng, Kexue Fu 0001, Zhijian Song, Yonghong Shi |
AAAI | 4 |
| 2026 | Semi-supervised medical image segmentation via anatomy-preserving consistency training
Shiman Li, Siqi Yin, Zhijian Song |
Image Vis. Comput. | 5 |
| 2026 | Diffusion-based controllable time-step consistency for semi-supervised medical image segmentation
Shiman Li, Jiayue Zhao, Chenxi Zhang 0004, Zhijian Song |
Pattern Recognit. | 5 |
| 2026 | DH-Mamba: Exploring Dual-Domain Hierarchical State Space Models for MRI Reconstruction
Yucong Meng, Kexue Fu 0001, Zhijian Song, Yonghong Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictions. Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced to generate CAMs in WSSS. However, previous WSSS methods solely adopt CLIP's vision-language paired property for dense localization, neglecting its inherently limited dense knowledge across both visual and text modalities, which renders CAM generation suboptimal. In this work, we propose DiCLIP, a novel WSSS framework that leverages the generative diffusion model to enhance CLIP's dense knowledge across two modalities. Specifically, Visual Correlation Enhancement (VCE) and Text Semantic Augmentation (TSA) modules are proposed for dense prediction enhancement. To improve the spatial awareness of visual features, our VCE module utilizes diffusion's reliable spatial consistency to mitigate the over-smoothing issue in CLIP's attention. It designs the Attention Clustering Refinement (ACR) module to reliably extract diverse correlation maps from the diffusion model. The correlation maps act as a diversity bias for CLIP's self-attention, recursively pushing its visual features towards a more discriminative dense distribution. To augment the semantics of text embeddings, our TSA module argues that a single text modality is insufficient to encompass the variability of visual categories. Thus, we leverage diffusion's generative power to maintain a dynamic key-value cache model, shifting CAM generation from a patch-text matching mechanism to a novel visual knowledge retrieval paradigm. With these enhancements, DiCLIP not only outperforms state-of-the-art methods on PASCAL VOC and MS COCO but also significantly reduces training costs. Code is publicly available at https://github.com/zwyang6/DiCLIP. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
IEEE Trans. Image Process. | 6 |
| 2026 | Deep Mutual Learning Among Partially Labeled Datasets for Multi-Organ SegmentationabstractLabeling multiple organs for segmentation is a complex and time-consuming process, resulting in a scarcity of comprehensively labeled multi-organ datasets while the emergence of numerous partially labeled datasets. Current methods face three critical limitations: incomplete exploitation of available supervision; complex inference, and insufficient validation of generalization capabilities. This paper proposes a new framework based on mutual learning, aiming to improve multi-organ segmentation performance by complementing information among partially labeled datasets. Specifically, this method consists of three key components: 1) partial-organ segmentation models training with Difference Mutual Learning, 2) pseudo-label generation and filtering, and 3) full-organ segmentation models training enhanced by Similarity Mutual Learning. Difference Mutual Learning enables each partial-organ segmentation model to utilize labels and features from other datasets as complementary signals, improving cross-dataset organ detection for better pseudo labels. Similarity Mutual Learning augments each full-organ segmentation model training with two additional supervision sources: inter-dataset ground truths and dynamic reliable transferred features, significantly boosting segmentation accuracy. The model obtained by this method achieves both high accuracy and efficient inference for multi-organ segmentation. Extensive experiments conducted on nine datasets spanning the head-neck, chest, abdomen, and pelvis demonstrate that the proposed method achieves SOTA performance. Linhao Qu, Ziyue Xie, Yonghong Shi, Zhijian Song |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Boosting ViT-based MRI Reconstruction from the Perspectives of Frequency Modulation, Spatial Purification, and Scale DiversificationabstractThe accelerated MRI reconstruction process presents a challenging ill-posed inverse problem due to the extensive under-sampling in k-space. Recently, Vision Transformers (ViTs) have become the mainstream for this task, demonstrating substantial performance improvements. However, there are still three significant issues remain unaddressed: (1) ViTs struggle to capture high-frequency components of images, limiting their ability to detect local textures and edge information, thereby impeding MRI restoration; (2) Previous methods calculate multi-head self-attention (MSA) among both related and unrelated tokens in content, introducing noise and significantly increasing computational burden; (3) The naive feed-forward network in ViTs cannot model the multi-scale information that is important for image restoration. In this paper, we propose FPS-Former, a powerful ViT-based framework, to address these issues from the perspectives of frequency modulation, spatial purification, and scale diversification. Specifically, for issue (1), we introduce a frequency modulation attention module to enhance the self-attention map by adaptively re-calibrating the frequency information in a Laplacian pyramid. For issue (2), we customize a spatial purification attention module to capture interactions among closely related tokens, thereby reducing redundant or irrelevant feature representations. For issue (3), we propose an efficient feed-forward network based on a hybrid-scale fusion strategy. Comprehensive experiments conducted on three public datasets show that our FPS-Former outperforms state-of-the-art methods while requiring lower computational costs. Yucong Meng, Yonghong Shi, Zhijian Song |
AAAI | 4 |
| 2025 | MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels typically uses Class Activation Maps (CAM) to achieve dense predictions. Recently, Vision Transformer (ViT) has provided an alternative to generate localization maps from class-patch attention. However, due to insufficient constraints on modeling such attention, we observe that the Localization Attention Maps (LAM) often struggle with the artifact issue, i.e., patch regions with minimal semantic relevance are falsely activated by class tokens. In this work, we propose MoRe to address this issue and further explore the potential of LAM. Our findings suggest that imposing additional regularization on class-patch attention is necessary. To this end, we first view the attention as a novel directed graph and propose the Graph Category Representation module to implicitly regularize the interaction among class-patch entities. It ensures that class tokens dynamically condense the related patch information and suppress unrelated artifacts at a graph level. Second, motivated by the observation that CAM from classification weights maintains smooth localization of objects, we devise the Localization-informed Regularization module to explicitly regularize the class-patch attention. It directly mines the token relations from CAM and further supervises the consistency between class and patch tokens in a learnable manner. Extensive experiments on PASCAL VOC and MS COCO validate that MoRe effectively addresses the artifact issue and achieves state-of-the-art performance, surpassing recent single-stage and even multi-stage methods. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
AAAI | 5 |
| 2025 | Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced in WSSS. However, recent methods primarily focus on image-text alignment for CAM generation, while CLIP’s potential in patch-text alignment remains unexplored. In this work, we propose ExCEL to explore CLIP’s dense knowledge via a novel patch-text alignment paradigm for WSSS. Specifically, we propose Text Semantic Enrichment (TSE) and Visual Calibration (VC) modules to improve the dense alignment across both text and vision modalities. To make text embeddings semantically informative, our TSE module applies Large Language Models (LLMs) to build a dataset-wide knowledge base and enriches the text representations with an implicit attribute-hunting process. To mine fine-grained knowledge from visual features, our VC module first proposes Static Visual Calibration (SVC) to propagate fine-grained knowledge in a non-parametric manner. Then Learnable Visual Calibration (LVC) is further proposed to dynamically shift the frozen features towards distributions with diverse semantics. With these enhancements, ExCEL not only retains CLIP’s training-free advantages but also significantly outperforms other state-of-the-art methods with much less training cost on PASCAL VOC and MS COCO. Code is available at https://github.com/zwyang6/ExCEL. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
CVPR | 6 |
| 2025 | Vector-Quantization-Driven Active Learning for Efficient Multi-modal Medical Segmentation with Cross-Modal Assistance
Xiaofei Du 0002, Haoran Wang 0009, Manning Wang, Zhijian Song |
MICCAI (7) | 4 |
| 2025 | Knowledge-Guided Multi-scale Graph Mamba for Whole Slide Image Classification
Minghong Duan, Yingfan Ma, Manning Wang, Zhijian Song |
MICCAI (12) | 5 |
| 2025 | Beyond the Distribution: Perturbation Toward Domain Distribution Boundary for Strengthening Generalizable Semi-Supervised SegmentationabstractThe generalization capability of models during inference in medical imaging, especially across data from different centers, is crucial, particularly in the context of limited data availability. Semi-supervised domain generalization learning has explored the use of unlabeled data to help the model generalize to unseen domains. However, existing methods primarily focus on decoupling and fusing information within the source distribution, resulting in limited improvement to the generalization ability. In contrast, we extend the exploration beyond source distributions and propose domain distribution boundary perturbation consistency learning for semi-supervised domain generalization in segmentation. We first use normalizing flows to map features to a mixture of Gaussian distributions, enabling the model to capture complex and diverse feature distributions across domains. Then, we perturb the features towards the domain boundaries and apply consistency regularization to maintain the quality of pseudo-labels while encouraging exploration of out-of-distribution features. This exploration aids the model's adaptation to unseen domains and enhances generalization. Our method is simple and effective, outperforming existing state-of-the-art approaches on three public benchmark datasets and achieving optimal generalization performance. Shiman Li, Jiayue Zhao, Chenxi Zhang 0004, Zhijian Song |
IEEE Signal Process. Lett. | 5 |
| 2025 | Tackling Ambiguity From Perspectives of Uncertainty Inference and Affinity Diversification for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve dense predictions without laborious annotations. However, due to the ambiguous contexts and fuzzy regions, the performance of WSSS, particularly during the stages of generating Class Activation Maps (CAMs) and refining pseudo masks, is widely hindered by ambiguity. Despite this, this issue has received little attention in previous literature. In this work, we propose UniA, a unified single-staged WSSS framework, to efficiently tackle this issue from the perspectives of uncertainty inference and affinity diversification. When activating class objects, we argue that the false activation stems from the bias to ambiguous regions during the feature extraction. Therefore, we formulate a robust feature representation with a Gaussian distribution and introduce the uncertainty estimation to avoid the bias. A distribution loss is proposed to supervise the process, which effectively captures the ambiguity and models the complex dependencies among features. When refining pseudo labels, we observe that the affinity from the prevailing refinement methods intends to be overly similar among ambiguities. To this end, we design an affinity diversification module to promote diversity among semantics. A mutual complementing refinement is first proposed to statically rectify the ambiguous affinity with multiple inferred pseudo labels. Then a contrastive affinity loss is further designed to dynamically diversify the relations among unrelated semantics. It stably propagates the diversity into the feature representation and helps generate better pseudo masks. Extensive experiments are conducted on PASCAL VOC, MS COCO, and medical ACDC datasets, which validate the efficiency of UniA tackling ambiguity and its superiority over recent single-staged or even most multi-staged competitors. Code is publicly available athttps://github.com/zwyang6/UniA. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
IEEE Trans. Multim. | 5 |
| 2024 | Separate and Conquer: Decoupling Co-occurrence via Decomposition and Representation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve segmentation tasks with-out dense annotations. However, attributed to the frequent coupling of co-occurring objects and the limited supervision from image-level labels, the challenging co-occurrence problem is widely present and leads to false activation of objects in WSSS. In this work, we devise a ‘Separate and Conquer’ scheme SeCo to tackle this issue from di-mensions of image space and feature space. In the im-age space, we propose to ‘separate’ the co-occurring ob-jects with image decomposition by subdividing images into patches. Importantly, we assign each patch a category tag from Class Activation Maps (CAMs), which spatially helps remove the co-context bias and guide the subsequent rep-resentation. In the feature space, we propose to ‘conquer’ the false activation by enhancing semantic representation with multi-granularity knowledge contrast. To this end, a dual-teacher-single-student architecture is designed and tag-guided contrast is conducted, which guarantee the cor-rectness of knowledge and further facilitate the discrepancy among co-contexts. We streamline the multi-staged WSSS pipeline end-to-end and tackle this issue without external supervision. Extensive experiments are conducted, validating the efficiency of our method and the superiority over previous single-staged and even multi-staged competitors on PASCAL VOC and MS COCO. Code is available here. Kexue Fu 0001, Minghong Duan, Linhao Qu, Shuo Wang 0011, Zhijian Song |
CVPR | 6 |
| 2024 | Trans2Fuse: Empowering image fusion through self-supervised learning and multi-modal transformations via transformer networks
Linhao Qu, Shaolei Liu, Manning Wang, Shiman Li, Siqi Yin, Zhijian Song |
Expert Syst. Appl. | 6 |
| 2024 | A comprehensive survey on deep active learning in medical image analysis
Haoran Wang 0009, Qiuye Jin, Shiman Li, Manning Wang, Zhijian Song |
Medical Image Anal. | 6 |
| 2024 | Wavelet-based spectrum transfer with collaborative learning for unsupervised bidirectional cross-modality domain adaptation on medical image segmentation
Shaolei Liu, Linhao Qu, Siqi Yin, Manning Wang, Zhijian Song |
Neural Comput. Appl. | 5 |
| 2024 | Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Good Instance Classifier Is All You NeedabstractWeakly supervised whole slide image classification is usually formulated as a multiple instance learning (MIL) problem, where each slide is treated as a bag, and the patches cut out of it are treated as instances. Existing methods either train an instance classifier through pseudo-labeling or aggregate instance features into a bag feature through attention mechanisms and then train a bag classifier, where the attention scores can be used for instance-level classification. However, the pseudo instance labels constructed by the former usually contain a lot of noise, and the attention scores constructed by the latter are not accurate enough, both of which affect their performance. In this paper, we propose an instance-level MIL framework based on contrastive learning and prototype learning to effectively accomplish both instance classification and bag classification tasks. To this end, we propose an instance-level weakly supervised contrastive learning algorithm for the first time under the MIL setting to effectively learn instance feature representation. We also propose an accurate pseudo label generation method through prototype learning. We then develop a joint training strategy for weakly supervised contrastive learning, prototype learning, and instance classifier training. Extensive experiments and visualizations on four datasets demonstrate the powerful performance of our method. Codes will be available. Linhao Qu, Yingfan Ma, Xiaoyuan Luo, Qinhao Guo, Manning Wang, Zhijian Song |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Negative Instance Guided Self-Distillation Framework for Whole Slide Image AnalysisabstractHistopathology image classification is an important clinical task, and current deep learning-based whole-slide image (WSI) classification methods typically cut WSIs into small patches and cast the problem as multi-instance learning. The mainstream approach is to train a bag-level classifier, but their performance on both slide classification and positive patch localization is limited because the instance-level information is not fully explored. In this article, we propose a negative instance-guided, self-distillation framework to directly train an instance-level classifier end-to-end. Instead of depending only on the self-supervised training of the teacher and the student classifiers in a typical self-distillation framework, we input the true negative instances into the student classifier to guide the classifier to better distinguish positive and negative instances. In addition, we propose a prediction bank to constrain the distribution of pseudo instance labels generated by the teacher classifier to prevent the self-distillation from falling into the degeneration of classifying all instances as negative. We conduct extensive experiments and analysis on three publicly available pathological datasets: CAMELYON16, PANDA, and TCGA, as well as an in-house pathological dataset for cervical cancer lymph node metastasis prediction. The results show that our method outperforms existing methods by a large margin. Code will be publicly available. Xiaoyuan Luo, Linhao Qu, Qinhao Guo, Zhijian Song, Manning Wang |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Boosting Whole Slide Image Classification from the Perspectives of Distribution, Correlation and MagnificationabstractBag-based multiple instance learning (MIL) methods have become the mainstream for Whole Slide Image (WSI) classification. However, there are still three important issues that have not been fully addressed: (1) positive bags with a low positive instance ratio are prone to the influence of a large number of negative instances; (2) the correlation between local and global features of pathology images has not been fully modeled; and (3) there is a lack of effective information interaction between different magnifications. In this paper, we propose MILBooster, a powerful dual-scale multi-stage MIL framework to address these issues from the perspectives of distribution, correlation, and magnification. Specifically, to address issue (1), we propose a plug-and-play bag filter that effectively increases the positive instance ratio of positive bags. For issue (2), we propose a novel window-based Transformer architecture called PiceBlock to model the correlation between local and global features of pathology images. For issue (3), we propose a dual-branch architecture to process different magnifications and design an information interaction module called Scale Mixer for efficient information interaction between them. We conducted extensive experiments on four clinical WSI classification tasks using three datasets. MILBooster achieved new state-of-the-art performance on all these tasks. Codes will be available at https://github.com/miccaiif/MILBooster. Linhao Qu, Minghong Duan, Yingfan Ma, Shuo Wang 0011, Manning Wang, Zhijian Song |
ICCV | 7 |
| 2023 | OpenAL: An Efficient Deep Active Learning Framework for Open-Set Pathology Image Classification
Linhao Qu, Yingfan Ma, Manning Wang, Zhijian Song |
MICCAI (2) | 5 |
| 2023 | The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationabstractThis paper introduces the novel concept of few-shot weakly supervised learning for pathology Whole Slide Image (WSI) classification, denoted as FSWC. A solution is proposed based on prompt learning and the utilization of a large language model, GPT-4. Since a WSI is too large and needs to be divided into patches for processing, WSI classification is commonly approached as a Multiple Instance Learning (MIL) problem. In this context, each WSI is considered a bag, and the obtained patches are treated as instances. The objective of FSWC is to classify both bags and instances with only a limited number of labeled bags. Unlike conventional few-shot learning problems, FSWC poses additional challenges due to its weak bag labels within the MIL framework. Drawing inspiration from the recent achievements of vision-language models (V-L models) in downstream few-shot classification tasks, we propose a two-level prompt learning MIL framework tailored for pathology, incorporating language prior knowledge. Specifically, we leverage CLIP to extract instance features for each patch, and introduce a prompt-guided pooling strategy to aggregate these instance features into a bag feature. Subsequently, we employ a small number of labeled bags to facilitate few-shot prompt learning based on the bag features. Our approach incorporates the utilization of GPT-4 in a question-and-answer mode to obtain language prior knowledge at both the instance and bag levels, which are then integrated into the instance and bag level language prompts. Additionally, a learnable component of the language prompts is trained using the available few-shot labeled data. We conduct extensive experiments on three real WSI datasets encompassing breast cancer, lung cancer, and cervical cancer, demonstrating the notable performance of the proposed method in bag and instance classification. All codes will be made publicly accessible. Linhao Qu, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang, Zhijian Song |
NeurIPS | 5 |
| 2023 | Density-based one-shot active learning for image segmentation
Qiuye Jin, Shiman Li, Xiaofei Du 0002, Mingzhi Yuan, Manning Wang, Zhijian Song |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | AIM-MEF: Multi-exposure image fusion based on adaptive information mining in both spatial and frequency domains
Linhao Qu, Siqi Yin, Shaolei Liu, Manning Wang, Zhijian Song |
Expert Syst. Appl. | 6 |
| 2023 | A learnable self-supervised task for unsupervised domain adaptation on point cloud classification and segmentation
Shaolei Liu, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang, Zhijian Song |
Frontiers Comput. Sci. | 5 |
| 2023 | A Structure-Aware Framework of Unsupervised Cross-Modality Domain Adaptation via Frequency and Spatial Knowledge DistillationabstractUnsupervised domain adaptation (UDA) aims to train a model on a labeled source domain and adapt it to an unlabeled target domain. In medical image segmentation field, most existing UDA methods rely on adversarial learning to address the domain gap between different image modalities. However, this process is complicated and inefficient. In this paper, we propose a simple yet effective UDA method based on both frequency and spatial domain transfer under a multi-teacher distillation framework. In the frequency domain, we introduce non-subsampled contourlet transform for identifying domain-invariant and domain-variant frequency components (DIFs and DVFs) and replace the DVFs of the source domain images with those of the target domain images while keeping the DIFs unchanged to narrow the domain gap. In the spatial domain, we propose a batch momentum update-based histogram matching strategy to minimize the domain-variant image style bias. Additionally, we further propose a dual contrastive learning module at both image and pixel levels to learn structure-related information. Our proposed method outperforms state-of-the-art methods on two cross-modality medical image segmentation datasets (cardiac and abdominal). Codes are avaliable at https://github.com/slliuEric/FSUDA. Shaolei Liu, Siqi Yin, Linhao Qu, Manning Wang, Zhijian Song |
IEEE Trans. Medical Imaging | 5 |
| 2022 | TransMEF: A Transformer-Based Multi-Exposure Image Fusion Framework Using Self-Supervised Multi-Task LearningabstractIn this paper, we propose TransMEF, a transformer-based multi-exposure image fusion framework that uses self-supervised multi-task learning. The framework is based on an encoder-decoder network, which can be trained on large natural image datasets and does not require ground truth fusion images. We design three self-supervised reconstruction tasks according to the characteristics of multi-exposure images and conduct these tasks simultaneously using multi-task learning; through this process, the network can learn the characteristics of multi-exposure images and extract more generalized features. In addition, to compensate for the defect in establishing long-range dependencies in CNN-based architectures, we design an encoder that combines a CNN module with a transformer module. This combination enables the network to focus on both local and global information. We evaluated our method and compared it to 11 competitive traditional and deep learning-based methods on the latest released multi-exposure image fusion benchmark dataset, and our method achieved the best performance in both subjective and objective evaluations. Code will be available at https://github.com/miccaiif/TransMEF. Linhao Qu, Shaolei Liu, Manning Wang, Zhijian Song |
AAAI | 4 |
| 2022 | DGMIL: Distribution Guided Multiple Instance Learning for Whole Slide Image Classification
Linhao Qu, Xiaoyuan Luo, Shaolei Liu, Manning Wang, Zhijian Song |
MICCAI (2) | 5 |
| 2022 | Bi-directional Weakly Supervised Knowledge Distillation for Whole Slide Image ClassificationabstractComputer-aided pathology diagnosis based on the classification of Whole Slide Image (WSI) plays an important role in clinical practice, and it is often formulated as a weakly-supervised Multiple Instance Learning (MIL) problem. Existing methods solve this problem from either a bag classification or an instance classification perspective. In this paper, we propose an end-to-end weakly supervised knowledge distillation framework (WENO) for WSI classification, which integrates a bag classifier and an instance classifier in a knowledge distillation framework to mutually improve the performance of both classifiers. Specifically, an attention-based bag classifier is used as the teacher network, which is trained with weak bag labels, and an instance classifier is used as the student network, which is trained using the normalized attention scores obtained from the teacher network as soft pseudo labels for the instances in positive bags. An instance feature extractor is shared between the teacher and the student to further enhance the knowledge exchange between them. In addition, we propose a hard positive instance mining strategy based on the output of the student network to force the teacher network to keep mining hard positive instances. WENO is a plug-and-play framework that can be easily applied to any existing attention-based bag classification methods. Extensive experiments on five datasets demonstrate the efficiency of WENO. Code is available at https://github.com/miccaiif/WENO. Linhao Qu, Xiaoyuan Luo, Manning Wang, Zhijian Song |
NeurIPS | 4 |
| 2022 | Globally Optimal Linear Model Fitting with Unit-Norm Constraint
Yinlong Liu, Manning Wang, Guang Chen 0001, Alois C. Knoll, Zhijian Song |
Int. J. Comput. Vis. | 6 |
| 2022 | Cold-start active learning for image classification
Qiuye Jin, Mingzhi Yuan, Shiman Li, Haoran Wang 0009, Manning Wang, Zhijian Song |
Inf. Sci. | 6 |
| 2022 | One-shot active learning for image segmentation via contrastive learning and diversity-based sampling
Qiuye Jin, Mingzhi Yuan, Qin Qiao, Zhijian Song |
Knowl. Based Syst. | 4 |
| 2022 | Deep active learning models for imbalanced image classification
Qiuye Jin, Mingzhi Yuan, Haoran Wang 0009, Manning Wang, Zhijian Song |
Knowl. Based Syst. | 5 |
| 2022 | Wavelet-based self-supervised learning for multi-scene image fusion
Shaolei Liu, Linhao Qu, Qin Qiao, Manning Wang, Zhijian Song |
Neural Comput. Appl. | 5 |
| 2022 | Vertebrae Labeling via End-to-End Integral Regression Localization and Multi-Label Classification NetworkabstractAccurate identification and localization of the vertebrae in CT scans is a critical and standard pre-processing step for clinical spinal diagnosis and treatment. Existing methods are mainly based on the integration of multiple neural networks, and most of them use heatmaps to locate the vertebrae's centroid. However, the process of obtaining vertebrae's centroid coordinates using heatmaps is non-differentiable, so it is impossible to train the network to label the vertebrae directly. Therefore, for end-to-end differential training of vertebrae coordinates on CT scans, a robust and accurate automatic vertebral labeling algorithm is proposed in this study. First, a novel end-to-end integral regression localization and multi-label classification network is developed, which can capture multi-scale features and also utilize the residual module and skip connection to fuse the multi-level features. Second, to solve the problem that the process of finding coordinates is non-differentiable and the spatial structure of location being destroyed, an integral regression module is used in the localization network. It combines the advantages of heatmaps representation and direct regression coordinates to achieve end-to-end training and can be compatible with any key point detection methods of medical images based on heatmaps. Finally, multi-label classification of vertebrae is carried out to improve the identification rate, which uses bidirectional long short-term memory (Bi-LSTM) online to enhance the learning of long contextual information of vertebrae. The proposed method is evaluated on a challenging data set, and the results are significantly better than state-of-the-art methods (identification rate is 91.1% and the mean localization error is 2.2 mm). The method is evaluated on a new CT data set, and the results show that our method has good generalization. Chunli Qin, Demin Yao, Han Zhuang, Yonghong Shi, Zhijian Song |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2021 | WaveFuse: A Unified Unsupervised Framework for Image Fusion with Discrete Wavelet Transform
Shaolei Liu, Manning Wang, Zhijian Song |
ICONIP (4) | 3 |
| 2021 | A Hybrid Attention Ensemble Framework for Zonal Prostate Segmentation
Mingyan Qiu, Zhijian Song |
MICCAI (1) | 3 |
| 2021 | Practical globally optimal consensus maximization by Branch-and-bound based on interval arithmetic
Yinlong Liu, Xuechen Li 0002, Chen Wang 0025, Manning Wang, Zhijian Song |
Pattern Recognit. | 6 |
| 2020 | GORFLM: Globally Optimal Robust Fitting for Linear Model
Yinlong Liu, Xuechen Li 0002, Chen Wang 0025, Manning Wang, Zhijian Song |
Signal Process. Image Commun. | 6 |
| 2019 | 2D-3D Point Set Registration Based on Global Rotation SearchabstractSimultaneously determining the relative pose and correspondence between a set of 3D points and its 2D projection is a fundamental problem in computer vision, and the problem becomes more difficult when the point sets are contaminated by noise and outliers. Traditionally, this problem is solved by local optimization methods, which usually start from an initial guess of the pose and alternately optimize the pose and the correspondence. In this paper, we formulate the problem as optimizing the pose of the 3D points in the SE(3) space to make its 2D projection best align with the 2D point set, which is measured by the cardinality of the inlier set on the 2D projection plane. We propose four geometric bounds for the position of the projection of a 3D point on the 2D projection plane and solve the 2D-3D point set registration problem by combining a global optimal rotation search and a grid search of translation. Compared with existing global optimization approaches, the proposed method utilizes a different problem formulation and more efficiently searches the translation space, which improves the registration speed. Experiments with synthetic and real data showed that the proposed approach significantly outperformed state-of-the-art local and global methods. Yinlong Liu, Zhijian Song, Manning Wang |
IEEE Trans. Image Process. | 3 |
| 2018 | Efficient Global Point Cloud Registration by Matching Rotation Invariant Features Through Translation Search
Yinlong Liu, Chen Wang 0025, Zhijian Song, Manning Wang |
ECCV (12) | 3 |
| 2011 | Deformable Registration for Geometric Distortion Correction of Diffusion Tensor Imaging
Xufeng Yao, Zhijian Song |
CAIP (1) | 2 |
| 2009 | Automatic localization of the center of fiducial markers in 3D CT/MRI images for image-guided neurosurgery
Manning Wang, Zhijian Song |
Pattern Recognit. Lett. | 2 |