VLDB 2026 Research / reviewers in the wild / expert
Tao Zhou 0002
dblp:98/4450-2
· DBLP profile ↗
115ranked-venue papers
26as first author
81since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 14 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 44 · 6 first-author · 34 since 2021Artificial intelligence and machine learning · 38 · 11 first-author · 25 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional Channel-selective Semantic Interaction for Semi-Supervised Medical SegmentationabstractSemi-supervised medical image segmentation is an effective method for addressing scenarios with limited labeled data. Existing methods mainly rely on frameworks such as mean teacher and dual-stream consistency learning. These approaches often face issues like error accumulation and model structural complexity, while also neglecting the interaction between labeled and unlabeled data streams. To overcome these challenges, we propose a Bidirectional Channel-selective Semantic Interaction (BCSI) framework for semi-supervised medical image segmentation. First, we propose a Semantic-Spatial Perturbation (SSP) mechanism, which disturbs the data using two strong augmentation operations and leverages unsupervised learning with pseudo-labels from weak augmentations. Additionally, we employ consistency on the predictions from the two strong augmentations to further improve model stability and robustness. Second, to reduce noise during the interaction between labeled and unlabeled data, we propose a Channel-selective Router (CR) component, which dynamically selects the most relevant channels for information exchange. This mechanism ensures that only highly relevant features are activated, minimizing unnecessary interference. Finally, the Bidirectional Channel-wise Interaction (BCI) strategy is employed to supplement additional semantic information and enhance the representation of important channels. Experimental results on multiple benchmarking 3D medical datasets demonstrate that the proposed method outperforms existing semi-supervised approaches. Kaiwen Huang 0002, Yizhe Zhang 0001, Yi Zhou 0007, Tianyang Xu 0001, Tao Zhou 0002 |
AAAI | 5 |
| 2026 | Cross-sample Consistency Learning for Semi-supervised Medical Image Segmentation
Tao Zhou 0002, Yunqi Gu, Kaiwen Huang 0002, Huazhu Fu, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 1 |
| 2026 | EvaNet: Toward More Efficient and Consistent Infrared and Visible Image Fusion AssessmentabstractEvaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide-and-conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no-reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning-based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks. Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Tao Zhou 0002, Hui Li 0037, Zhangyong Tang, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | A Theoretical Perspective on Streaming Noisy Data With Distribution ShiftabstractIntelligent systems typically need to continually learn from streaming data subject to distribution shift, where a key requirement is that they cannot catastrophically forget the historical knowledge learned from previous data. More seriously, streaming data often contain substantial label noise, which can exacerbate catastrophic forgetting and lead to performance degradation on forthcoming data. To address these problems, Continual Noisy Label Learning (CNLL) has been proposed. However, existing CNLL methods still fall short of the ability in addressing catastrophic forgetting because they adopted heuristic strategies in handling label noise and did not explicitly characterize the distributional shift across time, which hinders effective knowledge transfer from historical data to new data. To tackle these challenges, we theoretically analyze the problem of learning from streaming noisy data with distribution shift and propose a unified framework called Continual Noisy Label Learning on Drifting Data Streams (CNLDD). Specifically, we theoretically explore, for the first time, the upper bound of cumulative generalization error for CNLL problem, which reveals three factors leading to forgetting, namely selection bias of buffered data, distribution shift, and label noise. To alleviate the selection bias of buffered data, we design a two-step buffer update strategy to narrow the distribution gap between the original historical data and the selected representative data in buffer. To address distribution shift, our CNLDD explicitly characterizes the distribution discrepancies between buffered data and incoming data, prioritizing historical data with minimal discrepancies to enhance knowledge transfer. To tackle noisy labels, CNLDD estimates the importance weight of each example with the instance-dependent noise transition matrix, thereby avoiding the data bias and knowledge forgetting arising from noisy labels. Empirically, due to the unified modeling of the aforementioned issues, our CNLDD achieves superior classification performance when compared with state-of-the-art CNLL methods on both synthetic and real-world datasets. Wenshui Luo, Shuo Chen 0003, Tao Zhou 0002, Chen Gong 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Vision Mamba-enhanced Multi-level Context Aggregation Network for polyp segmentation
Jiaye Chen, Tianyang Xu 0001, Rui Wang 0050, Xiaoning Song, Tao Zhou 0002 |
Pattern Recognit. | 6 |
| 2026 | Frequency-enhanced contextual conversion network for esophageal lesion segmentation
Ziqi Tang, Xiaotong Niu, Long Rong, Yizhe Zhang 0001, Yawei Bi, Nan Ru, Longsong Li, Ningli Chai, Tao Zhou 0002 |
Pattern Recognit. | 9 |
| 2026 | Structured prompt-guided knowledge injection for medical image segmentation
Kelei He, Yizhe Zhang 0001, Yi Zhou 0007, Tao Zhou 0002, Dong Liang 0001 |
Pattern Recognit. | 5 |
| 2026 | HCRT: Hybrid network with correlation-aware region transformer for breast tumor segmentation in DCE-MRI
Lei Zheng 0017, Yuzhong Zhang, Tao Zhou 0002, Lei Zhou 0003, Dinggang Shen |
Pattern Recognit. | 4 |
| 2026 | Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image SegmentationabstractCross-domain few-shot medical image segmentation (CD-FSMIS) offers a promising and data-efficient solution for medical applications where annotations are severely scarce and multimodal analysis is required. However, existing methods typically filter out domain-specific information to improve generalization, which inadvertently limits cross-domain performance and degrades source-domain accuracy. To address this, we present Contrastive Graph Modeling (C-Graph), a framework that leverages the structural consistency of medical images as a reliable domain-transferable prior. We represent image features as graphs, with pixels as nodes and semantic affinities as edges. A Structural Prior Graph (SPG) layer is proposed to capture and transfer target-category node dependencies and enable global structure modeling through explicit node interactions. Building upon SPG layers, we introduce a Subgraph Matching Decoding (SMD) mechanism that exploits semantic relations among nodes to guide prediction. Furthermore, we design a Confusion-minimizing Node Contrast (CNC) loss to mitigate node ambiguity and subgraph heterogeneity by contrastively enhancing node discriminability in the graph space. Our method significantly outperforms prior CD-FSMIS approaches across multiple cross-domain benchmarks, achieving state-of-the-art performance while simultaneously preserving strong segmentation accuracy on the source domain. Our code is available at https://github.com/primebo1/C-Graph. Yuntian Bo, Tao Zhou 0002, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2026 | Uncertainty-Guided Prototype Reliability Enhancement Network for Few-Shot Medical Image SegmentationabstractFew-Shot Learning (FSL) has garnered increasing attention for data-scarce scenarios, particularly in medical segmentation tasks where only a few labeled data points are available. Existing few-shot segmentation methods typically learn prototypes from support images and employ nearest-neighbor searching to segment query images. Despite notable progress, effectively learning prototypes for each class remains a challenging task to achieve promising results. In this paper, we propose an Uncertainty-guided Prototype Reliability Enhancement Network (UPRE-Net) for few-shot medical image segmentation. Specifically, we present a dual-support branch to maximize the extraction of information from support images through augmentation techniques. To enhance the reliability of prototypes, we propose an Uncertainty-guided Prototype Generation (UPG) module. Within the UPG module, we first extract both global and local prototypes for each class and then apply uncertainty measures to select the most informative prototypes. Additionally, to effectively combine the prediction results from the dual-support branch, we present a Reliable Dynamic Fusion (RDF) module. This module dynamically integrates the two prediction results to generate a more reliable output. Furthermore, we present an Uncertainty-induced Weighted Loss (UWL) to ensure that the model pays more attention to these regions with high uncertainty. Experiments on four benchmark medical image datasets demonstrate that our proposed model significantly outperforms state-of-the-art methods. The code will be released at https://github.com/taozh2017/UPRENet. Tao Zhou 0002, Kaiwen Huang 0002, Yi Zhou 0007, Haofeng Zhang 0001, Boqiang Fan, Huazhu Fu |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Asymmetric Performance Profiling Using Foundation Models: Quantifying Reliability and Expert Capability in Medical AIabstractIn safety-critical domains like medical imaging, where diagnostic errors have severe consequences, AI models must be evaluated beyond average accuracy. A trustworthy model must demonstrate two distinct virtues: high reliability on common, easy cases and high expert capability on challenging, ambiguous, or rare cases. Conventional aggregate metrics fail to distinguish between these, masking a model's fatal flaw-such as misclassifying an easy case-by rewarding its high volume of trivial successes. We present Hardness-Aware Model Evaluation (HaME), a framework that assesses models using an asymmet-ric cost-benefit analysis. HaME identifies challenging instances via foundation models, then evaluates the target model using novel metrics (HaPrecision, HaRecall, HaFt, HaAUC) and the Brittleness Gap (B-Gap). Our formulation uniquely penalizes “easy” errors far more than “hard” failures, while simultaneously rewarding “hard” successes. This shifts the evaluation from “av-erage performance” to “clinical trustworthiness.” Experiments on medical image classification (Dermatology, Pneumonia, Retinal OCT) and segmentation (Nuclei) reveal that HaME uncovers critical reliability gaps and expert-level specializations invisible to standard metrics. Mingzhi Xu, Tao Zhou 0002, Qiang Chen 0004, Shuo Wang 0011, Yizhe Zhang 0001 |
BIBM | 3 |
| 2025 | A Correlation Manifold Self-Attention Network for EEG DecodingabstractRiemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces to geometrically capture the spatiotemporal dependencies inherent in time series data, e.g., electroencephalography (EEG). Recent studies have highlighted the full-rank correlation matrix as an advantageous alternative to the covariance matrix for data representation, owing to its invariance to the scale of variables. Motivated by these advancements, we propose the Correlation Attention Network (CorAtt) tailored for full-rank correlation matrices and implement it under the permutation-invariant and computationally efficient Off-Log and Log-Scaled geometries, respectively. Extensive evaluations on three benchmarking EEG datasets provide substantial evidence for the effectiveness of our introduced CorAtt. The code and supplementary material can be found at https://github.com/ChenHu-ML/CorAtt. Rui Wang 0050, Xiaoning Song, Tao Zhou 0002, Xiaojun Wu 0001, Nicu Sebe, Ziheng Chen 0001 |
IJCAI | 4 |
| 2025 | Text-Driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation
Kaiwen Huang 0002, Yi Zhou 0007, Huazhu Fu, Yizhe Zhang 0001, Chen Gong 0002, Tao Zhou 0002 |
MICCAI (5) | 6 |
| 2025 | Unsupervised Quality Control and Enhancement of Polyp Segmentation in Colonoscopy Videos Using Spatiotemporal Consistency
Tao Zhou 0002, Shuo Wang 0011, Yizhe Zhang 0001 |
MICCAI (10) | 2 |
| 2025 | Hierarchical Spatio-Temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos
Jian Yang 0003, Yizhe Zhang 0001, Tao Zhou 0002 |
MICCAI (3) | 4 |
| 2025 | Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
Yuang Yao, Yi Zhou 0007, Tao Zhou 0002 |
MICCAI (5) | 4 |
| 2025 | Diffusion-Based Multi-modal MR Fusion for TOF-MRA Image Synthesis
Tianen Yu, Lei Xiang 0001, Tao Zhou 0002 |
MICCAI (16) | 4 |
| 2025 | Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain AdaptationabstractClass-Incremental Unsupervised Domain Adaptation (CI-UDA) aims to adapt a model from a labeled source domain to an unlabeled target domain, where the sets of potential target classes appearing at different time steps are disjoint and are subsets of the source classes. The key to solving this problem lies in avoiding catastrophic forgetting of knowledge about previous target classes during continuously mitigating the domain shift. Most previous works cumbersomely combine two technical components. On one hand, they need to store and utilize rehearsal target sample from previous time steps to avoid catastrophic forgetting; on the other hand, they perform alignment only between classes shared across domains at each time step. Consequently, the memory will continuously increase and the asymmetric alignment may inevitably result in knowledge forgetting. In this paper, we propose to mine and preserve domain-invariant and class-agnostic knowledge to facilitate the CI-UDA task. Specifically, via using CLIP, we extract the class-agnostic properties which we name as ''attribute''. In our framework, we learn a ''key-value'' pair to represent an attribute, where the key corresponds to the visual prototype and the value is the textual prompt. We maintain two attribute dictionaries, each corresponding to a different domain. Then we perform attribute alignment across domains to mitigate the domain shift, via encouraging visual attention consistency and prediction consistency. Through attribute modeling and cross-domain alignment, we effectively reduce catastrophic knowledge forgetting while mitigating the domain shift, in a rehearsal-free way. Experiments on three CI-UDA benchmarks demonstrate that our method outperforms previous state-of-the-art methods and effectively alleviates catastrophic forgetting. Code is available at https://github.com/RyunMi/VisTA. Kerun Mi, Guoliang Kang, Lin Zhao 0003, Tao Zhou 0002, Chen Gong 0002 |
ACM Multimedia | 5 |
| 2025 | Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and BenchmarkingabstractUnifying multiple multi-modal visual object tracking (MMVOT) tasks draws increasing attention due to the complementary nature of different modalities in building robust tracking systems. Existing practices mix all data sensor types in a single training procedure, structuring a parallel paradigm from the data-centric perspective and aiming for a global optimum on the joint distribution of the involved tasks. However, the absence of a unified benchmark where all types of data coexist forces evaluations on separated benchmarks, causing inconsistency between training and testing, thus leading to performance degradation. To address these issues, this work advances in two aspects: A unified benchmark, coined as UniBench300, is introduced to bridge the inconsistency by incorporating multiple task data, reducing inference passes from three to one and cutting time consumption by 27%. The unification process is reformulated in a serial format, progressively integrating new tasks. In this way, the performance degradation can be specified as knowledge forgetting of previous tasks, which naturally aligns with the philosophy of continual learning (CL), motivating further exploration of injecting CL into the unification process. Extensive experiments conducted on two baselines and four benchmarks demonstrate the significance of UniBench300 and the superiority of CL in supporting a stable unification process. Moreover, while conducting dedicated analyses, the performance degradation is found to be negatively correlated with network capacity. Additionally, modality discrepancies contribute to varying degradation levels across tasks (RGBT > RGBD > RGBE in MMVOT), offering valuable insights for future multi-modal vision research. Source codes and the proposed benchmark is available at https://github.com/Zhangyong-Tang/UniBench300. Zhangyong Tang, Tianyang Xu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Tao Zhou 0002, Xiaojun Wu 0001, Josef Kittler |
ACM Multimedia | 5 |
| 2025 | Geometry-Aware Self-attention Network with Adaptive Log-Euclidean Metric for EEG Decoding
Zihao Bi, Rui Wang 0050, Tao Zhou 0002, Xiaoning Song, Xiaojun Wu 0001 |
PRCV (4) | 4 |
| 2025 | Coherence-Based Segmentation Quality Evaluator Trained on a Large Collection of Annotated Medical Images
Ahjol Senbi, Fei Lyu 0004, Qing Li 0001, Yuhui Tao, Qiang Chen 0004, Chengyan Wang, Shuo Wang 0011, Tao Zhou 0002, Yizhe Zhang 0001 |
PRCV (13) | 10 |
| 2025 | WeakPolyp-SAM: Segment Anything Model-driven weakly-supervised polyp segmentation
Tao Zhou 0002, Yunqi Gu, Yi Zhou 0007, Yizhe Zhang 0001, Ye Wu 0001, Huazhu Fu |
Knowl. Based Syst. | 2 |
| 2025 | SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 37 |
| 2025 | Dynamic Multi-scale Feature Integration Network for unsupervised MR-CT synthesis
Jiuming Jiang, Tao Zhou 0002, Yizhe Zhang 0001, Bin Qiu, Li Zhang 0021 |
Neural Networks | 4 |
| 2025 | Dual-scale enhanced and cross-generative consistency learning for semi-supervised medical image segmentation
Yunqi Gu, Tao Zhou 0002, Yizhe Zhang 0001, Yi Zhou 0007, Kelei He, Chen Gong 0002, Huazhu Fu |
Pattern Recognit. | 2 |
| 2025 | Frequency-Aware Interaction Network for Ultrasound Image SegmentationabstractAccurate segmentation of medical ultrasound images is crucial for guiding treatment decisions and assessing intervention effectiveness. The challenge of segmenting lesions in ultrasound images arises from factors such as low contrast, high speckle noise, artifacts, and blurred boundaries. Furthermore, this complexity varies significantly among lesions in different cases. While methods based on Convolutional Neural Networks (CNNs) and Transformers have shown promising results in this field, each approach possesses distinct advantages and limitations. To address these challenges, we propose a novel Frequency-aware Interaction Network (FINet). At the core of our FINet lies the proposed Multi-scale Frequency-aware Self-attention (MFS) module, which effectively captures multi-scale feature information within the self-attention layer. This enables our network to model both local and global features, capitalizing on the strengths of both CNNs and Transformers. Additionally, a frequency-aware network is introduced to learn the interactions between spatial locations in the frequency domain to enhance detailed feature representation such as edges. Furthermore, we present a collaborative interactive decoder network, in which a Selective Feature Interaction (SFI) module is proposed to facilitate the semantic and boundary feature interaction, resulting in more precise segmentation outcomes. Experimental results on four medical ultrasound image datasets show the superiority of our FINet over other state-of-the-art segmentation methods. More importantly, our model achieves an excellent trade-off between performance and computational efficiency. Tao Zhou 0002, Yizhe Zhang 0001, Shangbing Gao, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Uncertainty-Aware Cross-Training for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning has gained considerable popularity in medical image segmentation tasks due to its capability to reduce reliance on expert-examined annotations. Several mean-teacher (MT) based semi-supervised methods utilize consistency regularization to effectively leverage valuable information from unlabeled data. However, these methods often heavily rely on the student model and overlook the potential impact of cognitive biases within the model. Furthermore, some methods employ co-training using pseudo-labels derived from different inputs, yet generating high-confidence pseudo-labels from perturbed inputs during training remains a significant challenge. In this paper, we propose an Uncertainty-aware Cross-training framework for semi-supervised medical image Segmentation (UC-Seg). Our UC-Seg framework incorporates two distinct subnets to effectively explore and leverage the correlation between them, thereby mitigating cognitive biases within the model. Specifically, we present a Cross-subnet Consistency Preservation (CCP) strategy to enhance feature representation capability and ensure feature consistency across the two subnets. This strategy enables each subnet to correct its own biases and learn shared semantics from both labeled and unlabeled data. Additionally, we propose an Uncertainty-aware Pseudo-label Generation (UPG) component that leverages segmentation results and corresponding uncertainty maps from both subnets to generate high-confidence pseudo-labels. We extensively evaluate the proposed UC-Seg on various medical image segmentation tasks involving different modality images, such as MRI, CT, ultrasound, colonoscopy, and so on. The results demonstrate that our method achieves superior segmentation accuracy and generalization performance compared to other state-of-the-art semi-supervised methods. Our code and segmentation maps will be released at https://github.com/taozh2017/UCSeg. Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Xiaojun Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Cross-Domain Few-Shot Medical Image Segmentation via Dynamic Semantic MatchingabstractCross-domain few-shot medical image segmentation (CDFSMIS) presents the fundamental challenge of segmenting novel anatomical or tissue structures on unfamiliar medical imaging domains with limited annotated data. In this paper, we conduct an in-depth investigation of CDFSMIS and identify two critical observations: 1) the conventional matching mechanisms from existing few-shot models are particularly vulnerable to discrepancies in local characteristics between different domains and 2) the semantic representations learned from source domains often lack robustness when generalizing to unfamiliar target domains. Motivated by these insights, we propose a novel Dynamic Semantic Matching (DSM) framework that addresses these challenges through a three-component approach. First, we design a support-query feature re-weighting (SFR) mechanism that leverages multilevel hidden features to suppress domain-specific contents. Second, we introduce a dynamic semantic information selection (DSIS) strategy that adaptively identifies and combines domain-robust channels to construct generalizable representations. Third, we develop a dual-perspective semantic center calculation method to address the inherent texture imbalance in medical images. Extensive experiments on four unfamiliar target domains (MS-CMR, PI-PMR, Chest-X-Ray and ISIC2018) demonstrate that our approach significantly outperforms state-of-the-art few-shot segmentation and cross-domain few-shot segmentation models, validating the effectiveness of DSM in simultaneously addressing domain generalization and semantic matching challenges in medical image segmentation. The source code is available at https://github.com/YazhouZhu19/DSM. Yazhou Zhu 0001, Tao Zhou 0002, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Dual Interspersion and Flexible Deployment for Few-Shot Medical Image SegmentationabstractAcquiring a large volume of annotated medical data is impractical due to time, financial, and legal constraints. Consequently, few-shot medical image segmentation is increasingly emerging as a prominent research direction. Nowadays, Medical scenarios pose two major challenges: 1) intra-class variation caused by diversity among support and query sets; 2) inter-class extreme imbalance resulting from background heterogeneity. However, existing prototypical networks struggle to tackle these obstacles effectively. To this end, we propose a Dual Interspersion and Flexible Deployment (DIFD) model. Drawing inspiration from military interspersion tactics, we design the dual Interspersion module to generate representative basis prototypes from support features. These basis prototypes are then deeply interacted with query features. Furthermore, we introduce a fusion factor to fuse and refine the basis prototypes. Ultimately, we seamlessly integrate and flexibly deploy the basis prototypes to facilitate correct matching between the query features and basis prototypes, thus conducive to improving the segmentation accuracy of the model. Extensive experiments on three publicly available medical image datasets demonstrate that our model significantly outshines other SoTAs (2.78% higher dice score on average across all datasets), achieving a new level of performance. The code is available at: https://github.com/zmcheng9/DIFD. Ziming Cheng, Yang Long 0001, Tao Zhou 0002, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Learnable Prompting SAM-Induced Knowledge Distillation for Semi-Supervised Medical Image SegmentationabstractThe limited availability of labeled data has driven advancements in semi-supervised learning for medical image segmentation. Modern large-scale models tailored for general segmentation, such as the Segment Anything Model (SAM), have revealed robust generalization capabilities. However, applying these models directly to medical image segmentation still exposes performance degradation. In this paper, we propose a learnable prompting SAM-induced Knowledge distillation framework (KnowSAM) for semi-supervised medical image segmentation. Firstly, we propose a Multi-view Co-training (MC) strategy that employs two distinct sub-networks to employ a co-teaching paradigm, resulting in more robust outcomes. Secondly, we present a Learnable Prompt Strategy (LPS) to dynamically produce dense prompts and integrate an adapter to fine-tune SAM specifically for medical image segmentation tasks. Moreover, we propose SAM-induced Knowledge Distillation (SKD) to transfer useful knowledge from SAM to two sub-networks, enabling them to learn from SAM's predictions and alleviate the effects of incorrect pseudo-labels during training. Notably, the predictions generated by our subnets are used to produce mask prompts for SAM, facilitating effective inter-module information exchange. Extensive experimental results on various medical segmentation tasks demonstrate that our model outperforms the state-of-the-art semi-supervised segmentation approaches. Crucially, our SAM distillation framework can be seamlessly integrated into other semi-supervised segmentation methods to enhance performance. The code will be released upon acceptance of this manuscript at https://github.com/taozh2017/KnowSAM. Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Chen Gong 0002, Dong Liang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | On-the-Fly Improving Segment Anything for Medical Image Segmentation Using Auxiliary Online LearningabstractThe current variants of the Segment Anything Model (SAM), which include the original SAM and Medical SAM, still lack the capability to produce sufficiently accurate segmentation for medical images. In medical imaging contexts, it is not uncommon for human experts to rectify segmentations of specific test samples after SAM generates its segmentation predictions. These rectifications typically entail manual or semi-manual corrections employing state-of-the-art annotation tools. Motivated by this process, we introduce a novel approach that leverages the advantages of online machine learning to enhance Segment Anything (SA) during test time. We employ rectified annotations to perform online learning, with the aim of improving the segmentation quality of SA on medical images. To ensure the effectiveness and efficiency of online learning when integrated with large-scale vision models like SAM, we propose a new method called Auxiliary Online Learning (AuxOL), which entails adaptive online-batch and adaptive segmentation fusion. Experiments conducted on eight datasets covering four medical imaging modalities validate the effectiveness of the proposed method. Our work proposes and validates a new, practical, and effective approach for enhancing SA on downstream segmentation tasks (e.g., medical image segmentation). The code is publicly available at https://sam-auxol.github.io/AuxOL/. Tao Zhou 0002, Weidi Xie, Shuo Wang 0011, Qi Dou 0001, Yizhe Zhang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Domain-Interactive Contrastive Learning and Prototype-Guided Self-Training for Cross-Domain Polyp SegmentationabstractAccurate polyp segmentation plays a critical role in the diagnosis and treatment of colorectal cancer from colonoscopy images. While deep learning-based polyp segmentation models have made significant progress, they often suffer from performance degradation when applied to unseen target domain datasets collected from different imaging devices. To address this challenge, unsupervised domain adaptation (UDA) methods have gained attention by leveraging labeled source data and unlabeled target data to reduce the domain gap. However, existing UDA methods primarily focus on capturing class-wise representations, neglecting domain-wise representations. Additionally, uncertainty in pseudo-labels could hinder the segmentation performance. To tackle these issues, we propose a novel Domain-interactive Contrastive Learning and Prototype-guided Self-training (DCL-PS) framework for cross-domain polyp segmentation. Specifically, domain-interactive contrastive learning (DCL) with a domain-mixed prototype updating strategy is proposed to discriminate class-wise feature representations across domains. Then, to enhance the feature extraction ability of the encoder, we present a contrastive learning-based cross-consistency training (CL-CCT) strategy, which is imposed on both the prototypes obtained by the outputs of the main decoder and perturbed auxiliary outputs. Furthermore, we propose a prototype-guided self-training (PS) strategy, which dynamically assigns a weight for each pixel during self-training, filtering out unreliable pixels and improving the quality of pseudo-labels. Experimental results demonstrate the superiority of DCL-PS in improving polyp segmentation performance in the target domain. The code is released at https://github.com/taozh2017/DCLPS. Ziru Lu, Yizhe Zhang 0001, Yi Zhou 0007, Ye Wu 0001, Tao Zhou 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | MedSegViG: Medical Image Segmentation with a Vision Graph Neural NetworkabstractMedical image segmentation is a crucial step toward automatic clinical diagnosis, which has received growing interest. Although some existing methods based on convolutional neural networks or transformers have achieved remarkable success in this task, they still show limitations in effectively modeling the relationships among different objects in images. In this paper, we propose a novel deep learning based model to address this issue by leveraging a vision graph neural network (ViG). Our model, MedSegViG, mainly consists of a hierarchical ViG encoder and a lightweight convolutional decoder. The hierarchical encoder extracts multi-level features from the image and captures the object relationships with graph neural networks. The lightweight decoder then fuses these features and generates the corresponding segmentation map. Extensive experiments are conducted on seven datasets for three typical medical image segmentation tasks: polyp segmentation, skin lesion segmentation, and retinal vessel segmentation. The results demonstrate the superiority of our MedSegViG over state-of-the-art models across various tasks and datasets. The code is released on https://github.com/Xinhong-Li/MedSegViG. Geng Chen 0001, Yuanfeng Wu, Junqing Yang, Tao Zhou 0002, Yi Zhou 0007, Wentao Zhu 0002 |
BIBM | 5 |
| 2024 | Memory-Assisted Sub-Prototype Mining for Universal Domain AdaptationabstractUniversal domain adaptation aims to align the classes and reduce the feature gap between the same category of the source and target domains. The target private category is set as the unknown class during the adaptation process, as it is not included in the source domain. However, most existing methods overlook the intra-class structure within a category, especially in cases where there exists significant concept shift between the samples belonging to the same category. When samples with large concept shift are forced to be pushed together, it may negatively affect the adaptation performance. Moreover, from the interpretability aspect, it is unreasonable to align visual features with significant differences, such as fighter jets and civil aircraft, into the same category. Unfortunately, due to such semantic ambiguity and annotation cost, categories are not always classified in detail, making it difficult for the model to perform precise adaptation. To address these issues, we propose a novel Memory-Assisted Sub-Prototype Mining (MemSPM) method that can learn the differences between samples belonging to the same category and mine sub-classes when there exists significant concept shift between them. By doing so, our model learns a more reasonable feature space that enhances the transferability and reflects the inherent differences among samples annotated as the same category. We evaluate the effectiveness of our MemSPM method over multiple scenarios, including UniDA, OSDA, and PDA. Our method achieves state-of-the-art performance on four benchmarks in most cases. Yuxiang Lai, Yi Zhou 0007, Xinghong Liu, Tao Zhou 0002 |
ICLR | 4 |
| 2024 | MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
Chenran Zhang, Jianle Zhang, Yi Zhou 0007, Tao Zhou 0002, Huazhu Fu |
MICCAI (1) | 5 |
| 2024 | Customized Relationship Graph Neural Network for Brain Disorder Identification
Zhengwang Xia, Tao Zhou 0002, Jianfeng Lu 0003 |
MICCAI (2) | 3 |
| 2024 | SimTxtSeg: Weakly-Supervised Medical Image Segmentation with Simple Text Cues
Tao Zhou 0002, Yi Zhou 0007, Geng Chen 0001 |
MICCAI (8) | 2 |
| 2024 | TextPolyp: Point-Supervised Polyp Segmentation with Text Cues
Yi Zhou 0007, Yizhe Zhang 0001, Ye Wu 0001, Tao Zhou 0002 |
MICCAI (11) | 5 |
| 2024 | Combining Segment Anything Model with Domain-Specific Knowledge for Semi-Supervised Learning in Medical Image Segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Ye Wu 0001, Pengfei Gu, Shuo Wang 0011 |
PRCV (14) | 2 |
| 2024 | Inferring brain causal and temporal-lag networks for recognizing abnormal patterns of dementia
Zhengwang Xia, Tao Zhou 0002, Saqib Mamoon, Jianfeng Lu 0003 |
Medical Image Anal. | 2 |
| 2024 | TestFit: A plug-and-play one-pass test time method for medical image segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Yuhui Tao, Shuo Wang 0011, Ye Wu 0001, Benyuan Liu, Pengfei Gu, Qiang Chen 0004, Danny Ziyi Chen |
Medical Image Anal. | 2 |
| 2024 | Cross-contrast mutual fusion network for joint MRI reconstruction and super-resolution
Tao Zhou 0002, Lei Xiang 0001, Ye Wu 0001 |
Pattern Recognit. | 2 |
| 2024 | Uncertainty-Aware Hierarchical Aggregation Network for Medical Image SegmentationabstractMedical image segmentation is an essential process to assist clinics with computer-aided diagnosis and treatment. Recently, a large amount of convolutional neural network (CNN)-based methods have been rapidly developed and achieved remarkable performances in several different medical image segmentation tasks. However, the same type of infected region or lesions often has a diversity of scales, making it a challenging task to achieve accurate medical image segmentation. In this paper, we present a novel Uncertainty-aware Hierarchical Aggregation Network, namely UHA-Net, for medical image segmentation, which can fully make utilization of cross-level and multi-scale features to handle scale variations. Specifically, we propose a hierarchical feature fusion (HFF) module to aggregate high-level features, which is used to produce a global map for the coarse localization of the segmented target. Then, we propose an uncertainty-induced cross-level fusion (UCF) module to fully fuse features from the adjacent levels, which can learn knowledge guidance to capture the contextual information from adjacent resolutions. Further, a scale aggregation module (SAM) is presented to learn multi-scale features by using different convolution kernels, to effectively deal with scale variations. At last, we formulate a unified framework to simultaneously fuse inter-layer convolutional features and learn the discriminability of multi-scale representations from the intra-layer features, leading to accurate segmentation results. We carry out experiments on three different medical image segmentation tasks, and the results demonstrate that our UHA-Net outperforms state-of-the-art segmentation methods. Our implementation code and segmentation maps will be publicly at https://github.com/taozh2017/UHANet. Tao Zhou 0002, Yi Zhou 0007, Geng Chen 0001, Jianbing Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Few-Shot Medical Image Segmentation via Generating Multiple Representative DescriptorsabstractAutomatic medical image segmentation has witnessed significant development with the success of large models on massive datasets. However, acquiring and annotating vast medical image datasets often proves to be impractical due to the time consumption, specialized expertise requirements, and compliance with patient privacy standards, etc. As a result, Few-shot Medical Image Segmentation (FSMIS) has become an increasingly compelling research direction. Conventional FSMIS methods usually learn prototypes from support images and apply nearest-neighbor searching to segment the query images. However, only a single prototype cannot well represent the distribution of each class, thus leading to restricted performance. To address this problem, we propose to Generate Multiple Representative Descriptors (GMRD), which can comprehensively represent the commonality within the corresponding class distribution. In addition, we design a Multiple Affinity Maps based Prediction (MAMP) module to fuse the multiple affinity maps generated by the aforementioned descriptors. Furthermore, to address intra-class variation and enhance the representativeness of descriptors, we introduce two novel losses. Notably, our model is structured as a dual-path design to achieve a balance between foreground and background differences in medical images. Extensive experiments on four publicly available medical image datasets demonstrate that our method outperforms the state-of-the-art methods, and the detailed analysis also verifies the effectiveness of our designed module. Ziming Cheng, Tong Xin 0002, Tao Zhou 0002, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | COSTA: A Multi-Center TOF-MRA Dataset and a Style Self-Consistency Network for Cerebrovascular SegmentationabstractTime-of-flight magnetic resonance angiography (TOF-MRA) is the least invasive and ionizing radiation-free approach for cerebrovascular imaging, but variations in imaging artifacts across different clinical centers and imaging vendors result in inter-site and inter-vendor heterogeneity, making its accurate and robust cerebrovascular segmentation challenging. Moreover, the limited availability and quality of annotated data pose further challenges for segmentation methods to generalize well to unseen datasets. In this paper, we construct the largest and most diverse TOF-MRA dataset (COSTA) from 8 individual imaging centers, with all the volumes manually annotated. Then we propose a novel network for cerebrovascular segmentation, namely CESAR, with the ability to tackle feature granularity and image style heterogeneity issues. Specifically, a coarse-to-fine architecture is implemented to refine cerebrovascular segmentation in an iterative manner. An automatic feature selection module is proposed to selectively fuse global long-range dependencies and local contextual information of cerebrovascular structures. A style self-consistency loss is then introduced to explicitly align diverse styles of TOF-MRA images to a standardized one. Extensive experimental results on the COSTA dataset demonstrate the effectiveness of our CESAR network against state-of-the-art methods. We have made 6 subsets of COSTA with the source code online available, in order to promote relevant research in the community. Lei Mou, Jinghui Lin, Yifan Zhao 0001, Yonghuai Liu, Shaodong Ma, Jiong Zhang 0004, Wenhao Lv, Tao Zhou 0002, Jiang Liu 0001, Alejandro F. Frangi, Yitian Zhao |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Fusion-Embedding Siamese Network for Light Field Salient Object DetectionabstractLight field salient object detection (SOD) has shown remarkable success and gained considerable attention from the computer vision community. Existing methods usually employ a single-/two-stream network to detect saliency. However, these methods can only handle up to two different modalities at a time, preventing them from being able to fully explore the rich information in multi-modal light field derived data. To address this, we propose the first joint multi-modal learning framework, called FES-Net, for light field SOD, which can take rich inputs not limited to two modalities. Specifically, we propose an attention-aware adaptation module to first transform the multi-modal inputs for use in our joint learning framework. The transformed inputs are then fed to a Siamese network along with multiple embedded feature fusion modules to extract informative multi-modal features. Finally, we predict saliency maps from the high-level extracted features using a saliency decoder module. Our joint multi-modal learning framework effectively resolves the limitations of existing methods, providing efficient and effective multi-modal learning that can fully explore the valuable information in light field data for accurate saliency detection. Furthermore, we improve the performance by introducing the Transformer as our backbone network. To the best of our knowledge, the improved version of our model, called FES-Trans, is the first attempt to address the challenging light field SOD with the powerful Transformer technique. Extensive experiments on benchmark datasets demonstrate that our models are superior light field SOD approaches and outperform cutting-edge models remarkably. Geng Chen 0001, Huazhu Fu, Tao Zhou 0002, Guobao Xiao, Keren Fu, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Class-Wise Contrastive Prototype Learning for Semi-Supervised Classification Under Intersectional Class MismatchabstractTraditional Semi-Supervised Learning (SSL) classification methods focus on leveraging unlabeled data to improve the model performance under the setting where labeled set and unlabeled set share the same classes. Nevertheless, the above-mentioned setting is often inconsistent with many real-world circumstances. Practically, both the labeled set and unlabeled set often hold some individual classes, leading to an intersectional class-mismatch setting for SSL. Under this setting, existing SSL methods are often subject to performance degradation attributed to these individual classes. To solve the problem, we propose a Class-wise Contrastive Prototype Learning (CCPL) framework, which can properly utilize the unlabeled data to improve the SSL classification performance. Specifically, we employ a supervised prototype learning strategy and a class-wise contrastive separation strategy to construct a prototype for each known class. To reduce the influence of the individual classes in unlabeled set (i.e., out-of-distribution classes), each unlabeled example can be weighted reasonably based on the prototypes during classifier training, which helps to weaken the negative influence caused by out-of-distribution classes. To reduce the influence of the individual classes in labeled set (i.e., private classes), we present a private assignment suppression strategy to suppress the improper assignments of unlabeled examples to the private classes with the help of the prototypes. Experimental results on four benchmarks and one real-world dataset show that our CCPL has a clear advantage over fourteen representative SSL methods as well as two supervised learning methods under the intersectional class-mismatch setting. Tao Zhou 0002, Bo Han 0003, Tongliang Liu, Xinkai Liang, Chen Gong 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | Dynamic Weighted Adversarial Learning for Semi-Supervised Classification under Intersectional Class MismatchabstractNowadays, class-mismatch problem has drawn intensive attention in Semi-Supervised Learning (SSL), where the classes of labeled data are assumed to be only a subset of the classes of unlabeled data. However, in a more realistic scenario, the labeled data and unlabeled data often share some common classes while they also have their individual classes, which leads to an “intersectional class-mismatch” problem. As a result, existing SSL methods are often confused by these individual classes and suffer from performance degradation. To address this problem, we propose a novel Dynamic Weighted Adversarial Learning (DWAL) framework to properly utilize unlabeled data for boosting the SSL performance. Specifically, to handle the influence of the individual classes in unlabeled data (i.e., Out-Of-Distribution classes), we propose an enhanced adversarial domain adaptation to dynamically assign weight for each unlabeled example from the perspectives of domain adaptation and a class-wise weighting mechanism, which consists of transferability score and prediction confidence value. Besides, to handle the influence of the individual classes in labeled data (i.e., private classes), we propose a dissimilarity maximization strategy to suppress the inaccurate correlations caused by the examples of individual classes within labeled data. Therefore, our DWAL can properly make use of unlabeled data to acquire an accurate SSL classifier under intersectional class-mismatch setting, and extensive experimental results on five public datasets demonstrate the effectiveness of the proposed model over other state-of-the-art SSL methods. Tao Zhou 0002, Jian Yang 0003, Jie Yang 0002, Chen Gong 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Specificity-preserving RGB-D saliency detectionabstractRGB-D saliency detection has attracted increasing attention, due to its effectiveness and the fact that depth cues can now be conveniently captured. Existing works often focus on learning a shared representation through various fusion strategies, with few methods explicitly considering how to preserve modality-specific characteristics. In this paper, taking a new perspective, we propose a specificity-preserving network (SP-Net) for RGB-D saliency detection, which benefits saliency detection performance by exploring both the shared information and modality-specific properties (e.g., specificity). Specifically, two modality-specific networks and a shared learning network are adopted to generate individual and shared saliency maps. A cross-enhanced integration module (CIM) is proposed to fuse cross-modal features in the shared learning network, which are then propagated to the next layer for integrating cross-level information. Besides, we propose a multi-modal feature aggregation (MFA) module to integrate the modality-specific features from each individual decoder into the shared decoder, which can provide rich complementary multi-modal information to boost the saliency detection performance. Further, a skip connection is used to combine hierarchical features between the encoder and decoder layers. Experiments on six benchmark datasets demonstrate that our SP-Net outperforms other state-of-the-art methods. Code is available at: https://github.com/taozh2017/SPNet. Tao Zhou 0002, Deng-Ping Fan, Geng Chen 0001, Yi Zhou 0007, Huazhu Fu |
Comput. Vis. Media | 1 |
| 2023 | Combating medical noisy labels by disentangled distribution learning and consistency regularization
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Hanshi Sun |
Future Gener. Comput. Syst. | 3 |
| 2023 | Consistency and Diversity Induced Human Motion SegmentationabstractSubspace clustering is a classical technique that has been widely used for human motion segmentation and other related tasks. However, existing segmentation methods often cluster data without guidance from prior knowledge, resulting in unsatisfactory segmentation results. To this end, we propose a novel Consistency and Diversity induced human Motion Segmentation (CDMS) algorithm. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces, in which transfer subspace learning is conducted on different layers to capture multi-level information. A multi-mutual consistency learning strategy is carried out to reduce the domain gap between the source and target data. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Besides, a novel constraint based on the Hilbert Schmidt Independence Criterion (HSIC) is introduced to ensure the diversity of multi-level subspace representations, which enables the complementarity of multi-level representations to be explored to boost the transfer learning performance. Moreover, to preserve the temporal correlations, an enhanced graph regularizer is imposed on the learned representation coefficients and the multi-level representations of the source data. The proposed model can be efficiently solved using the Alternating Direction Method of Multipliers (ADMM) algorithm. Extensive experimental results on public human motion datasets demonstrate the effectiveness of our method against several state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Ling Shao 0001, Fatih Porikli, Haibin Ling, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Cross-level Feature Aggregation Network for Polyp Segmentation
Tao Zhou 0002, Yi Zhou 0007, Kelei He, Chen Gong 0002, Jian Yang 0003, Huazhu Fu, Dinggang Shen |
Pattern Recognit. | 1 |
| 2023 | Zero-Shot Camouflaged Object DetectionabstractThe goal of Camouflaged object detection (COD) is to detect objects that are visually embedded in their surroundings. Existing COD methods only focus on detecting camouflaged objects from seen classes, while they suffer from performance degradation to detect unseen classes. However, in a real-world scenario, collecting sufficient data for seen classes is extremely difficult and labeling them requires high professional skills, thereby making these COD methods not applicable. In this paper, we propose a new zero-shot COD framework (termed as ZSCOD), which can effectively detect the never unseen classes. Specifically, our framework includes a Dynamic Graph Searching Network (DGSNet) and a Camouflaged Visual Reasoning Generator (CVRG). In details, DGSNet is proposed to adaptively capture more edge details for boosting the COD performance. CVRG is utilized to produce pseudo-features that are closer to the real features of the seen camouflaged objects, which can transfer knowledge from seen classes to unseen classes to help detect unseen objects. Besides, our graph reasoning is built on a dynamic searching strategy, which can pay more attention to the boundaries of objects for reducing the influences of background. More importantly, we construct the first zero-shot COD benchmark based on the COD10K dataset. Experimental results on public datasets show that our ZSCOD not only detects the camouflaged object of unseen classes but also achieves state-of-the-art performance in detecting seen classes. Haoran Li 0024, Chun-Mei Feng 0001, Yong Xu 0001, Tao Zhou 0002, Lina Yao 0001, Xiaojun Chang |
IEEE Trans. Image Process. | 4 |
| 2023 | A Structure-Guided Effective and Temporal-Lag Connectivity Network for Revealing Brain Disorder MechanismsabstractBrain network provides important insights for the diagnosis of many brain disorders, and how to effectively model the brain structure has become one of the core issues in the domain of brain imaging analysis. Recently, various computational methods have been proposed to estimate the causal relationship (i.e., effective connectivity) between brain regions. Compared with traditional correlation-based methods, effective connectivity can provide the direction of information flow, which may provide additional information for the diagnosis of brain diseases. However, existing methods either ignore the fact that there is a temporal-lag in the information transmission across brain regions, or simply set the temporal-lag value between all brain regions to a fixed value. To overcome these issues, we design an effective temporal-lag neural network (termed ETLN) to simultaneously infer the causal relationships and the temporal-lag values between brain regions, which can be trained in an end-to-end manner. In addition, we also introduce three mechanisms to better guide the modeling of brain networks. The evaluation results on the Alzheimer's Disease Neuroimaging Initiative (ADNI) database demonstrate the effectiveness of the proposed method. Zhengwang Xia, Tao Zhou 0002, Saqib Mamoon, Amani Alfakih, Jianfeng Lu 0003 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Flexible Fusion Network for Multi-Modal Brain Tumor SegmentationabstractAutomated brain tumor segmentation is crucial for aiding brain disease diagnosis and evaluating disease progress. Currently, magnetic resonance imaging (MRI) is a routinely adopted approach in the field of brain tumor segmentation that can provide different modality images. It is critical to leverage multi-modal images to boost brain tumor segmentation performance. Existing works commonly concentrate on generating a shared representation by fusing multi-modal data, while few methods take into account modality-specific characteristics. Besides, how to efficiently fuse arbitrary numbers of modalities is still a difficult task. In this study, we present a flexible fusion network (termed F$^{2}$Net) for multi-modal brain tumor segmentation, which can flexibly fuse arbitrary numbers of multi-modal information to explore complementary information while maintaining the specific characteristics of each modality. Our F$^{2}$Net is based on the encoder-decoder structure, which utilizes two Transformer-based feature learning streams and a cross-modal shared learning network to extract individual and shared feature representations. To effectively integrate the knowledge from the multi-modality data, we propose a cross-modal feature-enhanced module (CFM) and a multi-modal collaboration module (MCM), which aims at fusing the multi-modal features into the shared learning network and incorporating the features from encoders into the shared decoder, respectively. Extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of our F$^{2}$Net over other state-of-the-art segmentation methods. Hengyi Yang, Tao Zhou 0002, Yi Zhou 0007, Yizhe Zhang 0001, Huazhu Fu |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Automatic Detection of Tooth-Gingiva Trim Lines on Dental SurfacesabstractDetecting the tooth-gingiva trim line from a dental surface plays a critical role in dental treatment planning and aligner 3D printing. Existing methods treat this task as a segmentation problem, which is resolved with geometric deep learning based mesh segmentation techniques. However, these methods can only provide indirect results (i.e., segmented teeth) and suffer from unsatisfactory accuracy due to the incapability of making full use of high-resolution dental surfaces. To this end, we propose a two-stage geometric deep learning framework for automatically detecting tooth-gingiva trim lines from dental surfaces. Our framework consists of a trim line proposal network (TLP-Net) for predicting an initial trim line from the low-resolution dental surface as well as a trim line refinement network (TLR-Net) for refining the initial trim line with the information from the high-resolution dental surface. Specifically, our TLP-Net predicts the initial trim line by fusing the multi-scale features from a U-Net with a proposed residual multi-scale attention fusion module. Moreover, we propose feature bridge modules and a trim line loss to further improve the accuracy. The resulting trim line is then fed to our TLR-Net, which is a deep-based LDDMM model with the high-resolution dental surface as input. In addition, dense connections are incorporated into TLR-Net for improved performance. Our framework provides an automatic solution to trim line detection by making full use of raw high-resolution dental surfaces. Extensive experiments on a clinical dental surface dataset demonstrate that our TLP-Net and TLR-Net are superior trim line detection methods and outperform cutting-edge methods in both qualitative and quantitative evaluations. Geng Chen 0001, Jie Qin 0004, Boulbaba Ben Amor, Weiming Zhou, Hang Dai, Tao Zhou 0002, Heyuan Huang, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Multi-Scale Transformer Network With Edge-Aware Pre-Training for Cross-Modality MR Image SynthesisabstractCross-modality magnetic resonance (MR) image synthesis can be used to generate missing modalities from given ones. Existing (supervised learning) methods often require a large number of paired multi-modal data to train an effective synthesis model. However, it is often challenging to obtain sufficient paired data for supervised training. In reality, we often have a small number of paired data while a large number of unpaired data. To take advantage of both paired and unpaired data, in this paper, we propose a Multi-scale Transformer Network (MT-Net) with edge-aware pre-training for cross-modality MR image synthesis. Specifically, an Edge-preserving Masked AutoEncoder (Edge-MAE) is first pre-trained in a self-supervised manner to simultaneously perform 1) image imputation for randomly masked patches in each image and 2) whole edge map estimation, which effectively learns both contextual and structural information. Besides, a novel patch-wise loss is proposed to enhance the performance of Edge-MAE by treating different masked patches differently according to the difficulties of their respective imputations. Based on this proposed pre-training, in the subsequent fine-tuning stage, a Dual-scale Selective Fusion (DSF) module is designed (in our MT-Net) to synthesize missing-modality images by integrating multi-scale features extracted from the encoder of the pre-trained Edge-MAE. Furthermore, this pre-trained encoder is also employed to extract high-level features from the synthesized image and corresponding ground-truth image, which are required to be similar (consistent) in the training. Experimental results show that our MT-Net achieves comparable performance to the competing methods even using 70% of all available paired data. Our code will be released at https://github.com/lyhkevin/MT-Net. Yonghao Li, Tao Zhou 0002, Kelei He, Yi Zhou 0007, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Automatic Schelling Point Detection From MeshesabstractMesh Schelling points explain how humans focus on specific regions of a 3D object. They have a large number of important applications in computer graphics and provide valuable information for perceptual psychology studies. However, detecting mesh Schelling points is time-consuming and expensive since the existing techniques are mostly based on participant observation studies. To overcome these limitations, we propose to employ powerful deep learning techniques to detect mesh Schelling points in an automatic manner, free from participant observation studies. Specifically, we utilize the mesh convolution and pooling operations to extract informative features from mesh objects, and then predict the 3D heat map of Schelling points in an end-to-end manner. In addition, we propose a Deep Schelling Network (DS-Net) to automatically detect the Schelling points, including a multi-scale fusion component and a novel region-specific loss function to improve our network for a better regression of heat maps. To the best of our knowledge, DS-Net is the first deep neural network for detecting Schelling points from 3D meshes. We evaluate DS-Net on a mesh Schelling point dataset obtained from participant observation studies. The experimental results demonstrate that DS-Net is capable of detecting mesh Schelling points effectively and outperforms various state-of-the-art mesh saliency methods and deep learning models, both qualitatively and quantitatively. Geng Chen 0001, Hang Dai, Tao Zhou 0002, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Delving into Local Features for Open-Set Domain Adaptation in Fundus Image Analysis
Yi Zhou 0007, Shaochen Bai, Tao Zhou 0002, Yu Zhang 0009, Huazhu Fu |
MICCAI (8) | 3 |
| 2022 | Multi-class Label Noise Learning via Loss Decomposition and Centroid EstimationabstractIn real-world scenarios, many large-scale datasets often contain inaccurate labels, i.e., noisy labels, which may confuse model training and lead to performance degradation. To overcome this issue, Label Noise Learning (LNL) has recently attracted much attention, and various methods have been proposed to design an unbiased risk estimator to the noise-free dataset to combat such label noise. Among them, a trend of works based on Loss Decomposition and Centroid Estimation (LDCE) has shown very promising performance. However, existing LNL methods based on LDCE are only designed for binary classification, and they are not directly extendable to multi-class situations. In this paper, we propose a novel multi-class robust learning method for LDCE, which is termed “MC-LDCE”. Specifically, we decompose the commonly adopted loss (e.g., mean squared loss) function into a label-dependent part and a label-independent part, in which only the former is influenced by label noise. Further, by defining a new form of data centroid, we transform the recovery problem of a label-dependent part to a centroid estimation problem. Finally, by critically examining the mathematical expectation of clean data centroid given the observed noisy set, the centroid can be estimated which helps to build an unbiased risk estimator for multi-class learning. The proposed MC-LDCE method is general and applicable to different types (i.e., linear and nonlinear) of classification models. The experimental results on five public datasets demonstrate the superiority of the proposed MC-LDCE against other representative LNL methods in tackling multi-class label noise problem. Yongliang Ding, Tao Zhou 0002, Yijing Luo, Chen Gong 0002 |
SDM | 2 |
| 2022 | Light field salient object detection: A review and benchmarkabstractSalient object detection (SOD) is a long-standing research topic in computer vision with increasing interest in the past decade. Since light fields record comprehensive information of natural scenes that benefit SOD in a number of ways, using light field inputs to improve saliency detection over conventional RGB inputs is an emerging trend. This paper provides the first comprehensive review and a benchmark for light field SOD, which has long been lacking in the saliency community. Firstly, we introduce light fields, including theory and data forms, and then review existing studies on light field SOD, covering ten traditional models, seven deep learning-based models, a comparative study, and a brief review. Existing datasets for light field SOD are also summarized. Secondly, we benchmark nine representative light field SOD models together with several cutting-edge RGB-D SOD models on four widely used light field datasets, providing insightful discussions and analyses, including a comparison between light field SOD and RGB-D SOD models. Due to the inconsistency of current datasets, we further generate complete data and supplement focal stacks, depth maps, and multi-view images for them, making them consistent and uniform. Our supplemental data make a universal benchmark possible. Lastly, light field SOD is a specialised problem, because of its diverse data representations and high dependency on acquisition hardware, so it differs greatly from other saliency detection tasks. We provide nine observations on challenges and future directions, and outline several open issues. All the materials including models, datasets, benchmarking results, and supplemented light field datasets are publicly available at https://github.com/kerenfu/LFSOD-Survey . Keren Fu, Yao Jiang 0002, Ge-Peng Ji, Tao Zhou 0002, Qijun Zhao, Deng-Ping Fan |
Comput. Vis. Media | 4 |
| 2022 | Three-dimensional affinity learning based multi-branch ensemble network for breast tumor segmentation in MRI
Lei Zhou 0003, Tao Zhou 0002, Fuhua Yan, Dinggang Shen |
Pattern Recognit. | 4 |
| 2022 | Camouflaged Object Detection via Context-Aware Cross-Level FusionabstractCamouflaged object detection (COD) aims to identify the objects that conceal themselves in natural scenes. Accurate COD suffers from a number of challenges associated with low boundary contrast and the large variation of object appearances, e.g., object size and shape. To address these challenges, we propose a novel Context-aware Cross-level Fusion Network ($\text{C}^{2}\text{F}$-Net), which fuses context-aware cross-level features for accurately identifying camouflaged objects. Specifically, we compute informative attention coefficients from multi-level features with our Attention-induced Cross-level Fusion Module (ACFM), which further integrates the features under the guidance of attention coefficients. We then propose a Dual-branch Global Context Module (DGCM) to refine the fused features for informative feature representations by exploiting rich global context information. Multiple ACFMs and DGCMs are integrated in a cascaded manner for generating a coarse prediction from high-level features. The coarse prediction acts as an attention map to refine the low-level features before passing them to our Camouflage Inference Module (CIM) to generate the final prediction. We perform extensive experiments on three widely used benchmark datasets and compare$\text{C}^{2}\text{F}$-Net with state-of-the-art (SOTA) models. The results show that$\text{C}^{2}\text{F}$-Net is an effective COD model and outperforms SOTA models remarkably. Further, an evaluation on polyp segmentation datasets demonstrates the promising potentials of our$\text{C}^{2}\text{F}$-Net in COD downstream applications. Our code is publicly available at:https://github.com/Ben57882/C2FNet-TSCVT Geng Chen 0001, Ge-Peng Ji, Ya-Feng Wu, Tao Zhou 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Marine Animal SegmentationabstractIn recent years, marine animal study has gained increasing research attention, which raises significant demands for fine-grained marine animal segmentation (MAS) techniques. In addition, deep learning has been widely adopted for object segmentation and has achieved promising performance. However, deep-based MAS is still lack of investigation due to the shortage of a large-scale MAS dataset. To tackle this issue, we construct the first large-scale MAS dataset, calledMAS3K, which consists of 3,103 images from different types, including camouflaged marine animal images, common marine animal images, and underwater images without marine animals. Furthermore, we consider different underwater conditions, such as low illumination, turbid water quality, photographic distortion, etc. Each image fromMAS3Kdataset has rich annotations, including an object-level mask, a category name, attributes, and a camouflage method (if applicable). Furthermore, we propose a novel MAS network, called Enhanced Cascade Decoder Network (ECD-Net), which consists of multiple Interactive Feature Enhancement Modules (IFEMs) and Cascade Decoder Modules (CDMs). InECD-Net, the IFEMs are first utilized to extract rich multi-scale features. The resulting features are then fed to the CDMs for accurately segmenting marine animals from complex underwater environments. We perform extensive experiments to compareECD-Netwith 10 cutting-edge object segmentation models. The results demonstrate thatECD-Netis an effective MAS model and outperforms the cutting-edge models, both qualitatively and quantitatively. Lin Li 0057, Bo Dong 0001, Eric Rigall, Tao Zhou 0002, Junyu Dong, Geng Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Dynamic Spectral-Spatial Poisson Learning for Hyperspectral Image Classification With Extremely Scarce LabelsabstractAcquiring labeled training examples for hyperspectral images (HSI) is an expensive task, and even labeling one more pixel requires a real-time field survey of tens of square meters. Therefore, it is highly demanded to achieve satisfactory accuracy for an HSI classification method when the number of labeled examples is extremely limited. However, most of the existing methods lack the ability to handle extremely sparse labeled data. To overcome this issue, we propose a novel graph-based framework for HSI classification, termed “dynamic spectral–spatial Poisson learning” (DSSPL). Specifically, three measures are used to enable the proposed model suitable for the situation of extremely limited labeled data. First, Poisson learning (PL) is adopted for predicting labels on a graph, as it can prevent undesirable constant output labels of traditional label propagation methods and generate more informative label determinations. Second, spectral and spatial graphs are constructed from various features and fused to build a spectral–spatial graph, which exploits comprehensive connective relationships among pixels. Third, in each iteration, the fused graph is dynamically updated by feeding back the up-to-date label information generated by each iteration. The feedback strategy progressively refines the fused graph, and the propagation on the updated graph in turn improves output labels iteratively. Intensive experimental results on three public datasets demonstrate that the proposed DSSPL significantly outperforms other state-of-the-art HSI classification methods when very few pixels (e.g., 3, 5, or 10 of each class) are labeled. Shengwei Zhong 0001, Tao Zhou 0002, Sheng Wan, Jian Yang 0003, Chen Gong 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Feature Aggregation and Propagation Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to detect/segment camouflaged objects embedded in the environment, which has attracted increasing attention over the past decades. Although several COD methods have been developed, they still suffer from unsatisfactory performance due to the intrinsic similarities between the foreground objects and background surroundings. In this paper, we propose a novel Feature Aggregation and Propagation Network (FAP-Net) for camouflaged object detection. Specifically, we propose a Boundary Guidance Module (BGM) to explicitly model the boundary characteristic, which can provide boundary-enhanced features to boost the COD performance. To capture the scale variations of the camouflaged objects, we propose a Multi-scale Feature Aggregation Module (MFAM) to characterize the multi-scale information from each layer and obtain the aggregated feature representations. Furthermore, we propose a Cross-level Fusion and Propagation Module (CFPM). In the CFPM, the feature fusion part can effectively integrate the features from adjacent layers to exploit the cross-level correlations, and the feature propagation part can transmit valuable context information from the encoder to the decoder network via a gate unit. Finally, we formulate a unified and end-to-end trainable framework where cross-level features can be effectively fused and propagated for capturing rich context information. Extensive experiments on three benchmark camouflaged datasets demonstrate that our FAP-Net outperforms other state-of-the-art COD models. Moreover, our model can be extended to the polyp segmentation task, and the comparison results further validate the effectiveness of the proposed model in segmenting polyps. The source code and results will be released at https://github.com/taozh2017/FAPNet. Tao Zhou 0002, Yi Zhou 0007, Chen Gong 0002, Jian Yang 0003, Yu Zhang 0009 |
IEEE Trans. Image Process. | 1 |
| 2022 | Self-Supervised Multi-Modal Hybrid Fusion Network for Brain Tumor SegmentationabstractAccurate medical image segmentation of brain tumors is necessary for the diagnosing, monitoring, and treating disease. In recent years, with the gradual emergence of multi-sequence magnetic resonance imaging (MRI), multi-modal MRI diagnosis has played an increasingly important role in the early diagnosis of brain tumors by providing complementary information for a given lesion. Different MRI modalities vary significantly in context, as well as in coarse and fine information. As the manual identification of brain tumors is very complicated, it usually requires the lengthy consultation of multiple experts. The automatic segmentation of brain tumors from MRI images can thus greatly reduce the workload of doctors and buy more time for treating patients. In this paper, we propose a multi-modal brain tumor segmentation framework that adopts the hybrid fusion of modality-specific features using a self-supervised learning strategy. The algorithm is based on a fully convolutional neural network. Firstly, we propose a multi-input architecture that learns independent features from multi-modal data, and can be adapted to different numbers of multi-modal inputs. Compared with single-modal multi-channel networks, our model provides a better feature extractor for segmentation tasks, which learns cross-modal information from multi-modal data. Secondly, we propose a new feature fusion scheme, named hybrid attentional fusion. This scheme enables the network to learn the hybrid representation of multiple features and capture the correlation information between them through an attention mechanism. Unlike popular methods, such as feature map concatenation, this scheme focuses on the complementarity between multi-modal data, which can significantly improve the segmentation results of specific regions. Thirdly, we propose a self-supervised learning strategy for brain tumor segmentation tasks. Our experimental results demonstrate the effectiveness of the proposed model against other state-of-the-art multi-modal medical segmentation methods. Feiyi Fang, Yazhou Yao, Tao Zhou 0002, Guosen Xie, Jianfeng Lu 0003 |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Guest Editorial Generative Adversarial Networks in Biomedical Image ComputingabstractThe papers in this special section focus on generative adversarial networks in biomedical image computing. The field of biomedical imaging has obtained great progress from Roentgen’s original discovery of the X-ray to the current imaging tools, including Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Computed Tomography (CT), and Ultrasound (US). The benefits of using these non-invasive imaging technologies are to assess the current condition of an organ or tissue, which can be used to monitor a patient over time over time for accurate and timely diagnosis and treatment.With the development of imaging technologies, developing advanced artificial intelligence algorithms for automated image analysis has shown the potential to change many aspects of clinical applications within the next decade. Meanwhile, these advanced technologies have also brought new issues and challenges. Thus, there has been a growing demand for biomedical imaging computing to be a component of clinical trials and device improvement. Currently, Generative adversarial networks (GANs) have been attached growing interests in the computer vision community due to their capability of data generation or translation. GAN-based models are able to learn from a set of training data and generate new data with the same characteristics as the training ones, which have also proven to be the state of the art for generating sharp and realistic images. More importantly, GAN has been rapidly applied to many traditional and novel applications in the medical domain, such as image reconstruction, segmentation, diagnosis, synthesis, and so on. Despite GAN substantial progress in these areas, their application to medical image computing still faces challenges and unsolved problems remain. Huazhu Fu, Tao Zhou 0002, Shuo Li 0001, Alejandro F. Frangi |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Doubly Supervised Transfer Classifier for Computer-Aided Diagnosis With Imbalanced ModalitiesabstractTransfer learning (TL) can effectively improve diagnosis accuracy of single-modal-imaging-based computer-aided diagnosis (CAD) by transferring knowledge from other related imaging modalities, which offers a way to alleviate the small-sample-size problem. However, medical imaging data generally have the following characteristics for the TL-based CAD: 1) The source domain generally has limited data, which increases the difficulty to explore transferable information for the target domain; 2) Samples in both domains often have been labeled for training the CAD model, but the existing TL methods cannot make full use of label information to improve knowledge transfer. In this work, we propose a novel doubly supervised transfer classifier (DSTC) algorithm. In particular, DSTC integrates the support vector machine plus (SVM+) classifier and the low-rank representation (LRR) into a unified framework. The former makes full use of the shared labels to guide the knowledge transfer between the paired data, while the latter adopts the block-diagonal low-rank (BLR) to perform supervised TL between the unpaired data. Furthermore, we introduce the Schatten-p norm for BLR to obtain a tighter approximation to the rank function. The proposed DSTC algorithm is evaluated on the Alzheimer's disease neuroimaging initiative (ADNI) dataset and the bimodal breast ultrasound image (BBUI) dataset. The experimental results verify the effectiveness of the proposed DSTC algorithm. Xiangmin Han, Xiaoyan Fei, Jun Wang 0024, Tao Zhou 0002, Shihui Ying, Jun Shi 0004, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Improving EEG Decoding via Clustering-Based Multitask Feature LearningabstractAccurate electroencephalogram (EEG) pattern decoding for specific mental tasks is one of the key steps for the development of brain-computer interface (BCI), which is quite challenging due to the considerably low signal-to-noise ratio of EEG collected at the brain scalp. Machine learning provides a promising technique to optimize EEG patterns toward better decoding accuracy. However, existing algorithms do not effectively explore the underlying data structure capturing the true EEG sample distribution and, hence, can only yield a suboptimal decoding accuracy. To uncover the intrinsic distribution structure of EEG data, we propose a clustering-based multitask feature learning algorithm for improved EEG pattern decoding. Specifically, we perform affinity propagation-based clustering to explore the subclasses (i.e., clusters) in each of the original classes and then assign each subclass a unique label based on a one-versus-all encoding strategy. With the encoded label matrix, we devise a novel multitask learning algorithm by exploiting the subclass relationship to jointly optimize the EEG pattern features from the uncovered subclasses. We then train a linear support vector machine with the optimized features for EEG pattern decoding. Extensive experimental studies are conducted on three EEG data sets to validate the effectiveness of our algorithm in comparison with other state-of-the-art approaches. The improved experimental results demonstrate the outstanding superiority of our algorithm, suggesting its prominent performance for EEG pattern decoding in BCI applications. Yu Zhang 0009, Tao Zhou 0002, Wei Wu 0022, Hua Xie, Hongru Zhu, Guoxu Zhou, Andrzej Cichocki |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Specificity-preserving RGB-D Saliency Detection
Tao Zhou 0002, Huazhu Fu, Geng Chen 0001, Yi Zhou 0007, Deng-Ping Fan, Ling Shao 0001 |
ICCV | 1 |
| 2021 | CCT-Net: Category-Invariant Cross-Domain Transfer for Medical Single-to-Multiple Disease DiagnosisabstractA medical imaging model is usually explored for the diagnosis of a single disease. However, with the expanding demand for multi-disease diagnosis in clinical applications, multi-function solutions need to be investigated. Previous works proposed to either exploit different disease labels to conduct transfer learning through fine-tuning, or transfer knowledge across different domains with similar diseases. However, these methods still cannot address the real clinical challenge - a multi-disease model is required but annotations for each disease are not always available. In this paper, we introduce the task of transferring knowledge from single-disease diagnosis (source domain) to enhance multi-disease diagnosis (target domain). A category-invariant cross-domain transfer (CCT) method is proposed to address this single-to-multiple extension. First, for domain-specific task learning, we present a confidence weighted pooling (CWP) to obtain coarse heatmaps for different disease categories. Then, conditioned on these heatmaps, category-invariant feature refinement (CIFR) blocks are proposed to better localize discriminative semantic regions related to the corresponding diseases. The category-invariant characteristic enables transferability from the source domain to the target domain. We validate our method in two popular areas: extending diabetic retinopathy to identifying multiple ocular diseases, and extending glioma identification to the diagnosis of other brain tumors. Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Ling Shao 0001 |
ICCV | 3 |
| 2021 | Visual-Textual Attentive Semantic Consistency for Medical Report GenerationabstractAutomatic report generation on medical radiographs have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding sizes, locations and other medical description patterns, which is essential for generating high-quality reports, is challenging. Although previous methods focused on producing readable reports, how to accurately detect and describe findings that match with the query X-Ray has not been successfully addressed. In this paper, we propose a multi-modality semantic attention model to integrate visual features, predicted key finding embeddings, as well as clinical features, and progressively decode reports with visual-textual semantic consistency. First, multi-modality features are extracted and attended with the hidden states from the sentence de-coder, to encode enriched context vectors for better decoding a report. These modalities include regional visual features of scans, semantic word embeddings of the top-K findings predicted with high probabilities, and clinical features of indications. Second, the progressive report decoder consists of a sentence decoder and a word decoder, where we propose image-sentence matching and description accuracy losses to constrain the visual-textual semantic consistency. Extensive experiments on the public MIMIC-CXR and IU X-Ray datasets show that our model achieves consistent improvements over the state-of-the-art methods. Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Huazhu Fu, Ling Shao 0001 |
ICCV | 3 |
| 2021 | Context-aware Cross-level Fusion Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) is a challenging task due to the low boundary contrast between the object and its surroundings. In addition, the appearance of camouflaged objects varies significantly, e.g., object size and shape, aggravating the difficulties of accurate COD. In this paper, we propose a novel Context-aware Cross-level Fusion Network (C2F-Net) to address the challenging COD task. Specifically, we propose an Attention-induced Cross-level Fusion Module (ACFM) to integrate the multi-level features with informative attention coefficients. The fused features are then fed to the proposed Dual-branch Global Context Module (DGCM), which yields multi-scale feature representations for exploiting rich global context information. In C2F-Net, the two modules are conducted on high-level features using a cascaded manner. Extensive experiments on three widely used benchmark datasets demonstrate that our C2F-Net is an effective COD model and outperforms state-of-the-art models remarkably. Our code is publicly available at: https://github.com/thograce/C2FNet. Geng Chen 0001, Tao Zhou 0002, Yi Zhang 0076, Nian Liu 0002 |
IJCAI | 3 |
| 2021 | Confidence-Aware Cascaded Network for Fetal Brain Segmentation on MR Images
Xukun Zhang, Zhiming Cui 0001, Changan Chen, Jingjiao Lou, Wenxin Hu, He Zhang 0023, Tao Zhou 0002, Feng Shi 0001, Dinggang Shen |
MICCAI (3) | 8 |
| 2021 | RGB-D salient object detection: A surveyabstractSalient object detection, which simulates human visual perception in locating the most significant object(s) in a scene, has been widely applied to various computer vision tasks. Now, the advent of depth sensors means that depth maps can easily be captured; this additional spatial information can boost the performance of salient object detection. Although various RGB-D based salient object detection models with promising performance have been proposed over the past several years, an in-depth understanding of these models and the challenges in this field remains lacking. In this paper, we provide a comprehensive survey of RGB-D based salient object detection models from various perspectives, and review related benchmark datasets in detail. Further, as light fields can also provide depth maps, we review salient object detection models and popular benchmark datasets from this domain too. Moreover, to investigate the ability of existing models to detect salient objects, we have carried out a comprehensive attribute-based evaluation of several representative RGB-D based salient object detection models. Finally, we discuss several challenges and open directions of RGB-D based salient object detection for future research. All collected models, benchmark datasets, datasets constructed for attribute-based evaluation, and related code are publicly available at https://github.com/taozh2017/RGBD-SODsurvey. Tao Zhou 0002, Deng-Ping Fan, Ming-Ming Cheng, Jianbing Shen, Ling Shao 0001 |
Comput. Vis. Media | 1 |
| 2021 | Multi-level cross-modal interaction network for RGB-D salient object detection
Zhou Huang 0001, Huai-Xin Chen, Tao Zhou 0002, Yun-Zhi Yang, Bi-Yuan Liu |
Neurocomputing | 3 |
| 2021 | Incomplete multi-modal representation learning for Alzheimer's disease diagnosis
Yanbei Liu, Lianxi Fan, Changqing Zhang 0002, Tao Zhou 0002, Zhitao Xiao, Lei Geng, Dinggang Shen |
Medical Image Anal. | 4 |
| 2021 | Momentum contrastive learning for few-shot COVID-19 diagnosis from chest CT images
Xiaocong Chen, Lina Yao 0001, Tao Zhou 0002, Jinming Dong, Yu Zhang 0009 |
Pattern Recognit. | 3 |
| 2021 | Contrast-weighted dictionary learning based saliency detection for VHR optical remote sensing images
Zhou Huang 0001, Huai-Xin Chen, Tao Zhou 0002, Yun-Zhi Yang, Chang-Yin Wang, Bi-Yuan Liu |
Pattern Recognit. | 3 |
| 2021 | Contrast-Attentive Thoracic Disease Recognition With Dual-Weighting Graph ReasoningabstractAutomatic thoracic disease diagnosis is a rising research topic in the medical imaging community, with many potential applications. However, the inconsistent appearances and high complexities of various lesions in chest X-rays currently hinder the development of a reliable and robust intelligent diagnosis system. Attending to the high-probability abnormal regions and exploiting the priori of a related knowledge graph offers one promising route to addressing these issues. As such, in this paper, we propose two contrastive abnormal attention models and a dual-weighting graph convolution to improve the performance of thoracic multi-disease recognition. First, a left-right lung contrastive network is designed to learn intra-attentive abnormal features to better identify the most common thoracic diseases, whose lesions rarely appear in both sides symmetrically. Moreover, an inter-contrastive abnormal attention model aims to compare the query scan with multiple anchor scans without lesions to compute the abnormal attention map. Once the intra- and inter-contrastive attentions are weighted over the features, in addition to the basic visual spatial convolution, a chest radiology graph is constructed for dual-weighting graph reasoning. Extensive experiments on the public NIH ChestX-ray and CheXpert datasets show that our model achieves consistent improvements over the state-of-the-art methods both on thoracic disease identification and localization. Yi Zhou 0007, Tianfei Zhou, Tao Zhou 0002, Huazhu Fu, Jiacheng Liu 0011, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Multi-Mutual Consistency Induced Transfer Subspace Learning for Human Motion SegmentationabstractHuman motion segmentation based on transfer subspace learning is a rising interest in action-related tasks. Although progress has been made, there are still several issues within the existing methods. First, existing methods transfer knowledge from source data to target tasks by learning domain-invariant features, but they ignore to preserve domain-specific knowledge. Second, the transfer subspace learning is employed in either low-level or high-level feature spaces, but few methods consider fusing multi-level features for subspace learning. To this end, we propose a novel multi-mutual consistency induced transfer subspace learning framework for human motion segmentation. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces and reduces the distribution gap between them through a multi-mutual consistency learning strategy. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Our model also conducts the transfer subspace learning on different layers to capture multi-level structural information. Further, to preserve the temporal correlations, we project the learned representations into a block-like space. The proposed model is efficiently optimized by using the Augmented Lagrange Multiplier (ALM) algorithm. Experimental results on four human motion datasets demonstrate the effectiveness of our method over other state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
CVPR | 1 |
| 2020 | PraNet: Parallel Reverse Attention Network for Polyp Segmentation
Deng-Ping Fan, Ge-Peng Ji, Tao Zhou 0002, Geng Chen 0001, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
MICCAI (6) | 3 |
| 2020 | M2 Net: Multi-modal Multi-channel Network for Overall Survival Time Prediction of Brain Tumor Patients
Tao Zhou 0002, Huazhu Fu, Yu Zhang 0009, Changqing Zhang 0002, Xiankai Lu, Jianbing Shen, Ling Shao 0001 |
MICCAI (2) | 1 |
| 2020 | Multi-modal latent space inducing ensemble SVM classifier for early dementia diagnosis with neuroimaging data
Tao Zhou 0002, Kim-Han Thung, Mingxia Liu 0001, Feng Shi 0001, Changqing Zhang 0002, Dinggang Shen |
Medical Image Anal. | 1 |
| 2020 | Synthesizing Talking Faces from Text and Audio: An Autoencoder and Sequence-to-Sequence Convolutional Neural Network
Na Liu 0007, Tao Zhou 0002, Yunfeng Ji, Lihong Wan |
Pattern Recognit. | 2 |
| 2020 | Multiview Latent Space Learning With Feature Redundancy MinimizationabstractMultiview learning has received extensive research interest and has demonstrated promising results in recent years. Despite the progress made, there are two significant challenges within multiview learning. First, some of the existing methods directly use original features to reconstruct data points without considering the issue of feature redundancy. Second, existing methods cannot fully exploit the complementary information across multiple views and meanwhile preserve the view-specific properties; therefore, the degraded learning performance will be generated. To address the above issues, we propose a novel multiview latent space learning framework with feature redundancy minimization. We aim to learn a latent space to mitigate the feature redundancy and use the learned representation to reconstruct every original data point. More specifically, we first project the original features from multiple views onto a latent space, and then learn a shared dictionary and view-specific dictionaries to, respectively, exploit the correlations across multiple views as well as preserve the view-specific properties. Furthermore, the Hilbert-Schmidt independence criterion is adopted as a diversity constraint to explore the complementarity of multiview representations, which further ensures the diversity from multiple views and preserves the local structure of the data in each view. Experimental results on six public datasets have demonstrated the effectiveness of our multiview learning approach against other state-of-the-art methods. Tao Zhou 0002, Changqing Zhang 0002, Chen Gong 0002, Harish Bhaskar, Jie Yang 0002 |
IEEE Trans. Cybern. | 1 |
| 2020 | Dual Shared-Specific Multiview Subspace ClusteringabstractMultiview subspace clustering has received significant attention as the availability of diverse of multidomain and multiview real-world data has rapidly increased in the recent years. Boosting the performance of multiview clustering algorithms is challenged by two major factors. First, since original features from multiview data are highly redundant, reconstruction based on these attributes inevitably results in inferior performance. Second, since each view of such multiview data may contain unique knowledge as against the others, it remains a challenge to exploit complimentary information across multiple views while simultaneously investigating the uniqueness of each view. In this paper, we present a novel dual shared-specific multiview subspace clustering (DSS-MSC) approach that simultaneously learns the correlations between shared information across multiple views and also utilizes view-specific information to depict specific property for each independent view. Further, we formulate a dual learning framework to capture shared-specific information into the dimensional reduction and self-representation processes, which strengthens the ability of our approach to exploit shared information while preserving view-specific property effectively. The experimental results on several benchmark datasets have demonstrated the effectiveness of the proposed approach against other state-of-the-art techniques. Tao Zhou 0002, Changqing Zhang 0002, Xi Peng 0001, Harish Bhaskar, Jie Yang 0002 |
IEEE Trans. Cybern. | 1 |
| 2020 | Inf-Net: Automatic COVID-19 Lung Infection Segmentation From CT ImagesabstractCoronavirus Disease 2019 (COVID-19) spread globally in early 2020, causing the world to face an existential health crisis. Automated detection of lung infections from computed tomography (CT) images offers a great potential to augment the traditional healthcare strategy for tackling COVID-19. However, segmenting infected regions from CT slices faces several challenges, including high variation in infection characteristics, and low intensity contrast between infections and normal tissues. Further, collecting a large amount of data is impractical within a short time period, inhibiting the training of a deep model. To address these challenges, a novel COVID-19 Lung Infection Segmentation Deep Network (Inf-Net) is proposed to automatically identify infected regions from chest CT slices. In our Inf-Net, a parallel partial decoder is used to aggregate the high-level features and generate a global map. Then, the implicit reverse attention and explicit edge-attention are utilized to model the boundaries and enhance the representations. Moreover, to alleviate the shortage of labeled data, we present a semi-supervised segmentation framework based on a randomly selected propagation strategy, which only requires a few labeled images and leverages primarily unlabeled data. Our semi-supervised framework can improve the learning ability and achieve a higher performance. Extensive experiments on our COVID-SemiSeg and real CT volumes demonstrate that the proposed Inf-Net outperforms most cutting-edge segmentation models and advances the state-of-the-art performance. Deng-Ping Fan, Tao Zhou 0002, Ge-Peng Ji, Yi Zhou 0007, Geng Chen 0001, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Hi-Net: Hybrid-Fusion Network for Multi-Modal MR Image SynthesisabstractMagnetic resonance imaging (MRI) is a widely used neuroimaging technique that can provide images of different contrasts (i.e., modalities). Fusing this multi-modal data has proven particularly effective for boosting model performance in many tasks. However, due to poor data quality and frequent patient dropout, collecting all modalities for every patient remains a challenge. Medical image synthesis has been proposed as an effective solution, where any missing modalities are synthesized from the existing ones. In this paper, we propose a novel Hybrid-fusion Network (Hi-Net) for multi-modal MR image synthesis, which learns a mapping from multi-modal source images (i.e., existing modalities) to target images (i.e., missing modalities). In our Hi-Net, a modality-specific network is utilized to learn representations for each individual modality, and a fusion network is employed to learn the common latent representation of multi-modal data. Then, a multi-modal synthesis network is designed to densely combine the latent representation with hierarchical features from each modality, acting as a generator to synthesize the target images. Moreover, a layer-wise multi-modal fusion strategy effectively exploits the correlations among multiple modalities, where a Mixed Fusion Block (MFB) is proposed to adaptively weight different fusion strategies. Extensive experiments demonstrate the proposed model outperforms other state-of-the-art medical image synthesis methods. Tao Zhou 0002, Huazhu Fu, Geng Chen 0001, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Inter-modality Dependence Induced Data Recovery for MCI Conversion Prediction
Tao Zhou 0002, Kim-Han Thung, Yu Zhang 0009, Huazhu Fu, Jianbing Shen, Dinggang Shen, Ling Shao 0001 |
MICCAI (4) | 1 |
| 2019 | Interpretable Feature Learning Using Multi-output Takagi-Sugeno-Kang Fuzzy System for Multi-center ASD Diagnosis
Jun Wang 0024, Tao Zhou 0002, Zhaohong Deng, Huifang Huang, Shitong Wang 0001, Jun Shi 0004, Dinggang Shen |
MICCAI (3) | 3 |
| 2019 | Deep Multi-modal Latent Representation Learning for Automated Dementia Diagnosis
Tao Zhou 0002, Mingxia Liu 0001, Huazhu Fu, Jun Wang 0024, Jianbing Shen, Ling Shao 0001, Dinggang Shen |
MICCAI (4) | 1 |
| 2019 | Latent Representation Learning for Alzheimer's Disease Diagnosis With Incomplete Multi-Modality Neuroimaging and Genetic DataabstractThe fusion of complementary information contained in multi-modality data [e.g., magnetic resonance imaging (MRI), positron emission tomography (PET), and genetic data] has advanced the progress of automated Alzheimer's disease (AD) diagnosis. However, multi-modality based AD diagnostic models are often hindered by the missing data, i.e., not all the subjects have complete multi-modality data. One simple solution used by many previous studies is to discard samples with missing modalities. However, this significantly reduces the number of training samples, thus leading to a sub-optimal classification model. Furthermore, when building the classification model, most existing methods simply concatenate features from different modalities into a single feature vector without considering their underlying associations. As features from different modalities are often closely related (e.g., MRI and PET features are extracted from the same brain region), utilizing their inter-modality associations may improve the robustness of the diagnostic model. To this end, we propose a novel latent representation learning method for multi-modality based AD diagnosis. Specifically, we use all the available samples (including samples with incomplete modality data) to learn a latent representation space. Within this space, we not only use samples with complete multi-modality data to learn a common latent representation, but also use samples with incomplete multi-modality data to learn independent modality-specific latent representations. We then project the latent representations to the label space for AD diagnosis. We perform experiments using 737 subjects from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database, and the experimental results verify the effectiveness of our proposed method. Tao Zhou 0002, Mingxia Liu 0001, Kim-Han Thung, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Multi-Layer Multi-View Classification for Alzheimer's Disease DiagnosisabstractIn this paper, we propose a novel multi-view learning method for Alzheimer's Disease (AD) diagnosis, using neuroimaging and genetics data. Generally, there are several major challenges associated with traditional classification methods on multi-source imaging and genetics data. First, the correlation between the extracted imaging features and class labels is generally complex, which often makes the traditional linear models ineffective. Second, medical data may be collected from different sources (i.e., multiple modalities of neuroimaging data, clinical scores or genetics measurements), therefore, how to effectively exploit the complementarity among multiple views is of great importance. In this paper, we propose a Multi-Layer Multi-View Classification (ML-MVC) approach, which regards the multi-view input as the first layer, and constructs a latent representation to explore the complex correlation between the features and class labels. This captures the high-order complementarity among different views, as we exploit the underlying information with a low-rank tensor regularization. Intrinsically, our formulation elegantly explores the nonlinear correlation together with complementarity among different views, and thus improves the accuracy of classification. Finally, the minimization problem is solved by the Alternating Direction Method of Multipliers (ADMM). Experimental results on Alzheimer's Disease Neuroimaging Initiative (ADNI) data sets validate the effectiveness of our proposed method. Changqing Zhang 0002, Ehsan Adeli-Mosabbeb, Tao Zhou 0002, Xiaobo Chen 0001, Dinggang Shen |
AAAI | 3 |
| 2018 | Synthesizing Missing PET from MRI with Cycle-consistent Generative Adversarial Networks for Alzheimer's Disease Diagnosis
Yongsheng Pan, Mingxia Liu 0001, Chunfeng Lian, Tao Zhou 0002, Yong Xia 0001, Dinggang Shen |
MICCAI (3) | 4 |
| 2018 | Online discriminative dictionary learning for robust object tracking
Tao Zhou 0002, Fanghui Liu 0001, Harish Bhaskar, Jie Yang 0002, Huanlong Zhang, Ping Cai |
Neurocomputing | 1 |
| 2018 | Inverse Nonnegative Local Coordinate Factorization for Visual TrackingabstractRecently, nonnegative matrix factorization (NMF) with part-based representation has been widely used for appearance modeling in visual tracking. Unfortunately, not all the targets can be successfully decomposed as “parts” unless some rigorous conditions are satisfied. To avoid this problem, this paper introduces NMF's variants into the visual tracking framework in the view of data clustering for appearance modeling. First, an initial target appearance model based on NMF is proposed to describe the target's appearance with the incorporated local coordinate factorization constraint, orthogonality of the bases, and L1,1norm regularized sparse residual error constraint. Second, an inverse NMF model is proposed in which each learned base vector is regarded as a clustering center in a low-dimensional subspace. Potential target samples (from the foreground) will be clustered around base vectors, while the candidate samples (from the background) are very likely to spread irregularly over the entire clustering space. Such differences can be fully exploited by the inverse NMF model to produce more discriminative encoding vectors than the conventional NMF method. Furthermore, incremental updating model is introduced into the tracking framework for online updating the initial appearance model. Experiments on object tracking benchmark suggest that our tracker is able to achieve promising performance when compared with some state-of-the-art methods in deformation, occlusion, and other challenging situations. Fanghui Liu 0001, Tao Zhou 0002, Chen Gong 0002, Keren Fu, Li Bai 0001, Jie Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Robust Visual Tracking via Online Discriminative and Low-Rank Dictionary LearningabstractIn this paper, we propose a novel and robust tracking framework based on online discriminative and low-rank dictionary learning. The primary aim of this paper is to obtain compact and low-rank dictionaries that can provide good discriminative representations of both target and background. We accomplish this by exploiting the recovery ability of low-rank matrices. That is if we assume that the data from the same class are linearly correlated, then the corresponding basis vectors learned from the training set of each class shall render the dictionary to become approximately low-rank. The proposed dictionary learning technique incorporates a reconstruction error that improves the reliability of classification. Also, a multiconstraint objective function is designed to enable active learning of a discriminative and robust dictionary. Further, an optimal solution is obtained by iteratively computing the dictionary, coefficients, and by simultaneously learning the classifier parameters. Finally, a simple yet effective likelihood function is implemented to estimate the optimal state of the target during tracking. Moreover, to make the dictionary adaptive to the variations of the target and background during tracking, an online update criterion is employed while learning the new dictionary. Experimental results on a publicly available benchmark dataset have demonstrated that the proposed tracking algorithm performs better than other state-of-the-art trackers. Tao Zhou 0002, Fanghui Liu 0001, Harish Bhaskar, Jie Yang 0002 |
IEEE Trans. Cybern. | 1 |
| 2018 | Robust Visual Tracking Revisited: From Correlation Filter to Template MatchingabstractIn this paper, we propose a novel matching based tracker by investigating the relationship between template matching and the recent popular correlation filter based trackers (CFTs). Compared to the correlation operation in CFTs, a sophisticated similarity metric termed mutual buddies similarity is proposed to exploit the relationship of multiple reciprocal nearest neighbors for target matching. By doing so, our tracker obtains powerful discriminative ability on distinguishing target and background as demonstrated by both empirical and theoretical analyses. Besides, instead of utilizing single template with the improper updating scheme in CFTs, we design a novel online template updating strategy named memory, which aims to select a certain amount of representative and reliable tracking results in history to construct the current stable and expressive template set. This scheme is beneficial for the proposed tracker to comprehensively understand the target appearance variations, recall some stable results. Both qualitative and quantitative evaluations on two benchmarks suggest that the proposed tracking method performs favorably against some recently developed CFTs and other competitive trackers. Fanghui Liu 0001, Chen Gong 0002, Xiaolin Huang, Tao Zhou 0002, Jie Yang 0002, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2017 | Visual tracking via structural patch-based dictionary pair learningabstractIn this paper, a novel visual tracking framework based on Structural Patch-based Dictionary Pair Learning (SPDPL), is proposed. The proposed representation model encapsulates partial and spatial structural variations of the target through a novel dictionary learning scheme. The proposed method facilitates learning a robust and discriminative dictionary by considering all patches from the same part of the target region as one class, thus transforming the tracking problem into a multi-class classification and reconstruction task. Finally, a simple yet effective observation model is designed to obtain the most optimal candidate during tracking. Systems experiments of the proposed tracking algorithm on the Object Tracker Benchmark (OTB) dataset have demonstrated improvements against several other state-of-the-art trackers. Tao Zhou 0002, Fanghui Liu 0001, Harish Bhaskar, Jie Yang 0002, Ping Cai |
ICIP | 1 |
| 2017 | Online learning and joint optimization of combined spatial-temporal models for robust visual tracking
Tao Zhou 0002, Harish Bhaskar, Fanghui Liu 0001, Jie Yang 0002, Ping Cai |
Neurocomputing | 1 |
| 2017 | Kernelized temporal locality learning for real-time visual tracking
Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Jie Yang 0002 |
Pattern Recognit. Lett. | 2 |
| 2017 | Graph Regularized and Locality-Constrained Coding for Robust Visual TrackingabstractVisual tracking is complicated due to factors, such as occlusion, background clutter, abrupt target motion, and illumination variations, among others. In recent years, subspace representation and sparse coding techniques have demonstrated significant improvements in tracking. However, performance gain in tracking has been at the expense of losing locality and similarity attributes among the instances to be encoded. In this paper, a graph regularized and locality-constrained coding (GRLC) technique that encapsulates local manifold structure of the data in order to preserve locality and similarity information among instances is proposed. The GRLC methodology incorporates a similarity-preserving term within the objective function of the locality-constrained linear coding model, thereby overcoming some of the inherent instability issues common to such coding methods. In the proposed GRLC scheme, a graph Laplacian regularizer is chosen as a smoothing operator to learn both the representation dictionary and the coefficients by preserving the local structure of the data. This graph Laplacian smoothing operator ensures that the representations vary smoothly along the geodesics of the data manifold. Thus, by deriving the objective function of the GRLC method, a discriminative dictionary of instances can be iteratively obtained and the corresponding coefficients for each candidate can be computed using this learned dictionary. Finally, an effective observation likelihood function based on reconstruction error and a simple dictionary update scheme for visual target tracking are also proposed. Experimental results on the CVPR2013 visual tracker benchmark have demonstrated a favorable performance of the proposed technique both in terms of accuracy and robustness. Tao Zhou 0002, Harish Bhaskar, Fanghui Liu 0001, Jie Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Visual Tracking via Nonnegative Multiple CodingabstractIt has been extensively observed that an accurate appearance model is critical to achieving satisfactory performance for robust object tracking. Most existing top-ranked methods rely on linear representation over a single dictionary, which brings about improper understanding on the target appearance. To address this problem, in this paper, we propose a novel appearance model named as “nonnegative multiple coding” (NMC) to accurately represent a target. First, a series of local dictionaries are created with different predefined numbers of nearest neighbors, and then the contributions of these dictionaries are automatically learned. As a result, this ensemble of dictionaries can comprehensively exploit the appearance information carried by all the constituted dictionaries. Second, the existing methods explicitly impose the nonnegative constraint to coefficient vectors, but in the proposed model, we directly deploy an efficient 12 norm regularization to achieve the similar nonnegative purpose with theoretical guarantees. Moreover, an efficient occlusion detection scheme is designed to alleviate tracking drifts, which investigates whether negative templates are selected to represent the severely occluded target. Experimental results on two benchmarks demonstrate that our NMC tracker are able to achieve superior performance to state-of-the-art methods. Fanghui Liu 0001, Chen Gong 0002, Tao Zhou 0002, Keren Fu, Xiangjian He, Jie Yang 0002 |
IEEE Trans. Multim. | 3 |
| 2016 | Robust visual tracking via inverse nonnegative matrix factorizationabstractThe establishment of robust target appearance model over time is an overriding concern in visual tracking. In this paper, we propose an inverse nonnegative matrix factorization (NMF) method for robust appearance modeling. Rather than using a linear combination of nonnegative basis vectors for each target image patch in conventional NMF, the proposed method is a reverse thought to conventional NMF tracker. It utilizes both the foreground and background information, and imposes a local coordinate constraint, where the basis matrix is sparse matrix from the linear combination of candidates with corresponding nonnegative coefficient vectors. Inverse NMF is used as a feature encoder, where the resulting coefficient vectors are fed into a SVM classifier for separating the target from the background. The proposed method is tested on several videos and compared with seven state-of-the-art methods. Our results have provided further support to the effectiveness and robustness of the proposed method. Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Irene Y. H. Gu, Jie Yang 0002 |
ICASSP | 2 |
| 2016 | Correlation filter tracking via bootstrap learningabstractIn recent years, correlation filter based trackers outperform better than other trackers. Nevertheless, they only employ one feature and a single kernel, so they are usually not robust in complex scenes. In this paper, we derive a multi-feature and multi-kernel correlation filter based tracker which fully takes advantage of the invariance-discriminative power spectrums of various features and kernels to further improve the performance. A novel bootstrap learning method is utilized to obtain a strong classifier by fusing these weak kernel correlation filters (KCFs). Moreover, a new target scale estimation strategy is incorporated into our framework. The efficient and effective scale estimation method is based on target dictionary representation. The proposed method is tested on several videos and compared with seven state-of-the-art methods. Experimental results have provided further support to the effectiveness and robustness of the proposed method. Kunqi Gu, Tao Zhou 0002, Fanghui Liu 0001, Jie Yang 0002, Yu Qiao 0003 |
ICIP | 2 |
| 2016 | Incremental Robust Nonnegative Matrix Factorization for Object Tracking
Fanghui Liu 0001, Mingna Liu, Tao Zhou 0002, Yu Qiao 0003, Jie Yang 0002 |
ICONIP (2) | 3 |
| 2016 | Geometric affine transformation estimation via correlation filter for visual tracking
Fanghui Liu 0001, Tao Zhou 0002, Jie Yang 0002 |
Neurocomputing | 2 |
| 2016 | Robust visual tracking via constrained correlation filter coding
Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Jie Yang 0002 |
Pattern Recognit. Lett. | 2 |
| 2015 | Small target detection using an optimization-based filterabstractSmall target detection is a critical problem in the Infrared Search And Track (IRST) system. Although it has been studied for years, there are some challenges remained, e.g. cloud edges and horizontal lines are likely to cause false alarms. This paper proposes a novel method using an optimization-based filter to detect infrared small target in heavy clutter. First, we design a certain pixel area as active area. Second, a weighted quadratic cost function is performed in the active area. Finally, a filter based on statistics of active area is derived from the cost function. Our method could preserve heterogeneous area, meanwhile, remove target region. Experimental results show our method achieves satisfied performance in heavy clutter. Keren Fu, Tao Zhou 0002, Jie Yang 0002, Qiang Wu 0001, Xiangjian He |
ICASSP | 3 |
| 2015 | Online learning of multi-feature weights for robust object trackingabstractSparse Representation based Classification (SRC) and its potential in object tracking have been explored in recent years. However, the trade-off between the discriminative ability of the overly emphasized sparse representation and the lack of insight on correlation of visual information has raised questions over the general applicability of such methods in object tracking. In addition, the need for the optimization of a series of l1-regularized least square norm, increases the computational complexity thereby limiting their usage in real-time applications. In this paper, a novel approach to robust object tracking is proposed. First, the variations in the appearance of the tracked target is modelled using PCA basis vectors, and further, a l2-regularized least square method is used to solve the proposed representation model. In order to improve the robustness of feature representation in object tracking applications, weights are associated with multiple trackers; each formulated using a different feature, and adapted via an online learning scheme. Finally, a decision fusion criterion is imposed to generate an optimized output through the weighted combination of different tracking results. Experiments on challenging video sequences have demonstrated the superior accuracy and robustness of the proposed method in comparison to thirteen other state-of-the-art baselines. Tao Zhou 0002, Harish Bhaskar, Jie Yang 0002, Xiangjian He |
ICIP | 1 |
| 2015 | Robust visual tracking via efficient manifold ranking with low-dimensional compressive features
Tao Zhou 0002, Xiangjian He, Keren Fu, Jie Yang 0002 |
Pattern Recognit. | 1 |
| 2014 | Visual tracking based on weighted subspace reconstruction errorabstractIt is a challenging task to develop an effective and robust visual tracking method due to factors such as pose variation, illumination change, occlusion, and motion blur. In this paper, a novel tracking algorithm based on weighted subspace reconstruction error is proposed. We first compute the discriminative weights by sparse construction error with template dictionary consisted of positive and negative samples, and then confidence map for candidates is computed through subspace reconstruction error. Finally, the location of the target object is estimated by maximizing the decision map which is combined discriminative weights and subspace reconstruction error. Furthermore, we use the new evaluation criterion to verify the robustness of the current tracking result, which can reduce the accumulated error effectively. Experimental results on some challenging video sequences show that the proposed algorithm performs favorably against seven state-of-the-art methods in terms of accuracy and robustness. Tao Zhou 0002, Jie Yang 0002, Xiangjian He |
ICIP | 1 |
| 2014 | Visual tracking via graph-based efficient manifold ranking with low-dimensional compressive featuresabstractIn this paper, a novel and robust tracking method based on efficient manifold ranking is proposed. For tracking, tracked results are taken as labeled nodes while candidate samples are taken as unlabeled nodes, and the goal of tracking is to search the unlabeled sample that is the most relevant with existing labeled nodes by manifold ranking algorithm. Meanwhile, we adopt non-adaptive random projections to preserve the structure of original image space, and a very sparse measurement matrix is used to efficiently extract low-dimensional compressive features for object representation. Furthermore, spatial context is used to improve the robustness to appearance variations. Experimental results on some challenging video sequences show the proposed algorithm outperforms six state-of-the-art methods in terms of accuracy and robustness. Tao Zhou 0002, Xiangjian He, Keren Fu, Jie Yang 0002 |
ICME | 1 |