EDBT 2026 Demo / reviewers in the wild / expert
Fengtao Zhou
dblp:281/9853
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
Sheng Huang 0001, Fengtao Zhou, Bo Liu 0005, Qingshan Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Scan-Invariant Mamba With Differentiated Sequence Contrastive Learning in Computational PathologyabstractMultiple instance learning (MIL) is a commonly used paradigm for histopathological analysis due to the ultra-high resolution and coarse-grained labels of Whole Slide Images (WSIs). Recent studies apply Mamba architecture to WSI classification by modeling MIL as long-sequence tasks, but a key discrepancy remains: Mamba's output is sensitive to scanning modes, whereas MIL requires scan-invariant predictions. To address this problem, we propose Scan-invariant Mamba with Differentiated Sequence Contrastive Learning (SMDC-MIL), a novel Mamba-based MIL approach enabling bag-level feature learning independent of input modes. Our method mitigates scanning-mode impacts and adapts Mamba to learn the bag discrimination features that are independent of the input mode via two innovations: 1) a differentiated sequence generation mechanism that employs instance rearrangement, augmentation, and masking to simulate real-world scanning variations by maximizing differences in sequence order, length, and composition from the same WSI; and 2) a differentiated sequence contrastive learning architecture that enforces consistent bag-level representations and predictions across diverse sequences using the same Mamba model, guiding it to prioritize scan-invariant discriminative features. Experimental results on 4 computational pathology tasks and 10 datasets demonstrate that our SMDC-MIL achieves state-of-the-art performance compared to other methods. The corresponding code is available at https://github.com/LianYueZ/SMDCMIL.git. Sheng Huang 0001, Xin Zhang 0131, Bo Liu 0005, Fengtao Zhou, Kang Li 0004, Hao Chen 0011, Meng Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Distilled Prompt Learning for Incomplete Multimodal Survival PredictionabstractThe integration of multimodal data including pathology images and gene profiles is widely applied in precise survival prediction. Despite recent advances in multimodal survival models, collecting complete modalities for multi-modal fusion still poses a significant challenge, hindering their application in clinical settings. Current approaches tackling incomplete modalities often fall short, as they typically compensate for only a limited part of the knowledge of missing modalities. To address this issue, we propose a Distilled Prompt Learning framework (DisPro) to utilize the strong robustness of Large Language Models (LLMs) to missing modalities, which employs two-stage prompting for compensation of comprehensive information for missing modalities. In the first stage, Unimodal Prompting (UniPro) distills the knowledge distribution of each modality, preparing for supplementing modality-specific knowledge of the missing modality in the subsequent stage. In the second stage, Multimodal Prompting (MultiPro) leverages available modalities as prompts for LLMs to infer the missing modality, which provides modality-common information. Simultaneously, the unimodal knowledge acquired in the first stage is injected into multimodal inference to compensate for the modality-specific knowledge of the missing modality. Extensive experiments covering various missing scenarios demonstrated the superiority of the proposed method. The code is available at https://github.com/Innse/DisPro. Yingxue Xu, Fengtao Zhou, Yihui Wang 0002, Hao Chen 0011 |
CVPR | 2 |
| 2025 | Revisiting End-to-End Learning with Slide-level Supervision in Computational PathologyabstractPre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in computational pathology (CPath), benefiting cancer diagnosis and prognosis.
However, performance limitations arise from the absence of encoder fine-tuning for downstream tasks and disjoint optimization with MIL. While slide-level supervised end-to-end (E2E) learning
is an intuitive solution to this issue, it faces challenges such as high computational demands and suboptimal results.
These limitations motivate us to revisit E2E learning.
We argue that prior work neglects inherent E2E optimization challenges, leading to performance disparities compared to traditional two-stage methods.
In this paper, we pioneer the elucidation of optimization challenge caused by sparse-attention MIL and propose a novel MIL called ABMILX.
ABMILX mitigates this problem through global correlation-based attention refinement and multi-head mechanisms.
With the efficient multi-scale random patch sampling strategy, an E2E trained ResNet with ABMILX surpasses SOTA foundation models under the two-stage paradigm across multiple challenging benchmarks,
while remaining computationally efficient ($<$ 10 RTX3090 GPU hours).
We demonstrate the potential of E2E learning in CPath and
calls for greater research focus in this area.
The code is https://github.com/DearCaat/E2E-WSI-ABMILX. Fengtao Zhou, Xiang Li 0041, Ming-Ming Cheng |
NeurIPS | 4 |
| 2025 | Cohort-Individual Cooperative Learning for Multimodal Cancer Survival AnalysisabstractRecently, we have witnessed impressive achievements in cancer survival analysis by integrating multimodal data, e.g., pathology images and genomic profiles. However, the heterogeneity and high dimensionality of these modalities pose significant challenges in extracting discriminative representations while maintaining good generalization. In this paper, we propose a Cohort-individual Cooperative Learning (CCL) framework to advance cancer survival analysis by collaborating knowledge decomposition and cohort guidance. Specifically, first, we propose a Multimodal Knowledge Decomposition (MKD) module to explicitly decompose multimodal knowledge into four distinct components: redundancy, synergy, and uniqueness of the two modalities. Such a comprehensive decomposition can enlighten the models to perceive easily overlooked yet important information, facilitating an effective multimodal fusion. Second, we propose a Cohort Guidance Modeling (CGM) to mitigate the risk of overfitting task-irrelevant information. It can promote a more comprehensive and robust understanding of the underlying multimodal data while avoiding the pitfalls of overfitting and enhancing the generalization ability of the model. By cooperating with the knowledge decomposition and cohort guidance methods, we develop a robust multimodal survival analysis model with enhanced discrimination and generalization abilities. Extensive experimental results on five cancer datasets demonstrate the effectiveness of our model in integrating multimodal data for survival analysis. Our code is available at https://github.com/moothes/CCL-survival. Huajun Zhou, Fengtao Zhou, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyabstractMultiple instance learning (MIL) is the most widely used framework in computational pathology, encompassing sub-typing, diagnosis, prognosis, and more. However, the ex-isting MIL paradigm typically requires an offline instance feature extractor, such as a pre-trained ResNet or a foun-dation model. This approach lacks the capability for feature fine-tuning within the specific downstream tasks, limiting its adaptability and performance. To address this issue, we propose a Re-embedded Regional Transformer (R2T) for re-embedding the instance features online, which captures fine-grained local features and establishes connections across different regions. Unlike existing works that focus on pre-training powerful feature extractor or designing sophisticated instance aggregator, R2T is tailored to re-embed instance features online. It serves as a portable module that can seamlessly integrate into mainstream MIL models. Extensive experimental results on common computational pathology tasks validate that: 1) feature re-embedding improves the performance of MIL models based on ResNet-50 features to the level of foundation model features, and further enhances the performance of foundation model features; 2) the R2T can introduce more signifi-cant performance improvements to various MIL models; 3) R2T-MIL, as an R2T-enhanced AB-MIL, outperforms other latest methods by a large margin. The code is available at: https://github.com/DearCaat/RRT-MIL. Fengtao Zhou, Sheng Huang 0001, Yi Zhang 0113, Bo Liu 0005 |
CVPR | 2 |
| 2024 | Explain via Any Concept: Concept Bottleneck Model with Open Vocabulary Concepts
Andong Tan, Fengtao Zhou, Hao Chen 0011 |
ECCV (86) | 2 |
| 2024 | Post-hoc Part-Prototype NetworksabstractPost-hoc explainability methods such as Grad-CAM are popular because they do not influence the performance of a trained model. However, they mainly reveal ”where” a model looks at for a given input, fail to explain ”what” the model looks for (e.g., what is important to classify a bird image to a Scott Oriole?). Existing part-prototype networks leverage part-prototypes (e.g., characteristic Scott Oriole’s wing and head) to answer both ”where" and ”what", but often under-perform their black box counterparts in the accuracy. Therefore, a natural question is: can one construct a network that answers both ”where” and ”what" in a post-hoc manner to guarantee the model’s performance? To this end, we propose the first post-hoc part-prototype network via decomposing the classification head of a trained model into a set of interpretable part-prototypes. Concretely, we propose an unsupervised prototype discovery and refining strategy to obtain prototypes that can precisely reconstruct the classification head, yet being interpretable. Besides guaranteeing the performance, we show that our network offers more faithful explanations qualitatively and yields even better part-prototypes quantitatively than prior part-prototype networks. Andong Tan, Fengtao Zhou, Hao Chen 0011 |
ICML | 2 |
| 2023 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image ClassificationabstractThe whole slide image (WSI) classification is often formulated as a multiple instance learning (MIL) problem. Since the positive tissue is only a small fraction of the gigapixel WSI, existing MIL methods intuitively focus on identifying salient instances via attention mechanisms. However, this leads to a bias towards easy-to-classify instances while neglecting hard-to-classify instances. Some literature has revealed that hard examples are beneficial for modeling a discriminative boundary accurately. By applying such an idea at the instance level, we elaborate a novel MIL framework with masked hard instance mining (MHIM-MIL), which uses a Siamese structure (Teacher-Student) with a consistency constraint to explore the potential hard instances. With several instance masking strategies based on attention scores, MHIM-MIL employs a momentum teacher to implicitly mine hard instances for training the student model, which can be any attention-based MIL model. This counter-intuitive strategy essentially enables the student to learn a better discriminating boundary. Moreover, the student is used to update the teacher with an exponential moving average (EMA), which in turn identifies new hard instances for subsequent training iterations and stabilizes the optimization. Experimental results on the CAMELYON-16 and TCGA Lung Cancer datasets demonstrate that MHIM-MIL outperforms other latest methods in terms of performance and training cost. The code is available at: https://github.com/DearCaat/MHIM-MIL. Sheng Huang 0001, Xiaoxian Zhang, Fengtao Zhou, Yi Zhang 0113, Bo Liu 0005 |
ICCV | 4 |
| 2023 | Cross-Modal Translation and Alignment for Survival AnalysisabstractWith the rapid advances in high-throughput sequencing technologies, the focus of survival analysis has shifted from examining clinical indicators to incorporating genomic profiles with pathological images. However, existing methods either directly adopt a straightforward fusion of pathological features and genomic profiles for survival prediction, or take genomic profiles as guidance to integrate the features of pathological images. The former would overlook intrinsic cross-modal correlations. The latter would discard pathological information irrelevant to gene expression. To address these issues, we present a Cross-Modal Translation and Alignment (CMTA) framework to explore the intrinsic cross-modal correlations and transfer potential complementary information. Specifically, we construct two parallel encoder-decoder structures for multi-modal data to integrate intra-modal information and generate cross-modal representation. Taking the generated cross-modal representation to enhance and recalibrate intra-modal representation can significantly improve its discrimination for comprehensive survival analysis. To explore the intrinsic cross-modal correlations, we further design a cross-modal attention module as the information bridge between different modalities to perform cross-modal interactions and transfer complementary information. Our extensive experiments on five public TCGA datasets demonstrate that our proposed framework outperforms the state-of-the-art methods. The source code has been released†. Fengtao Zhou, Hao Chen 0011 |
ICCV | 1 |
| 2022 | Boosting Multi-Label Image Classification with Complementary Parallel Self-DistillationabstractMulti-Label Image Classification (MLIC) appro-aches usually exploit label correlations to achieve good performance. However, emphasizing correlation like co-occurrence may overlook discriminative features and lead to model overfitting. In this study, we propose a generic framework named Parallel Self-Distillation (PSD) for boosting MLIC models. PSD decomposes the original MLIC task into several simpler MLIC sub-tasks via two elaborated complementary task decomposition strategies named Co-occurrence Graph Partition (CGP) and Dis-occurrence Graph Partition (DGP). Then, the MLIC models of fewer categories are trained with these sub-tasks in parallel for respectively learning the joint patterns and the category-specific patterns of labels. Finally, knowledge distillation is leveraged to learn a compact global ensemble of full categories with these learned patterns for reconciling the label correlation exploitation and model overfitting. Extensive results on MS-COCO and NUS-WIDE datasets demonstrate that our framework can be easily plugged into many MLIC approaches and improve performances of recent state-of-the-art approaches. The source code is released at https://github.com/Robbie-Xu/CPSD. Jiazhi Xu, Sheng Huang 0001, Fengtao Zhou, Luwen Huangfu, Daniel Dajun Zeng, Bo Liu 0005 |
IJCAI | 3 |
| 2022 | Multi-Label Image Classification via Category Prototype Compositional LearningabstractReal-world images are often compositions of multiple objects with different categories, scales, poses and locations. Adding nonexistent objects to an image (composing) or removing existent objects from an image (decomposing) leads to higher discrepancy in appearance, which reveals an important but long-neglected compositional nature of multi-label images. In light of this observation, we propose a novel end-to-end compositional learning framework named Category Prototype Compositional Learning (CPCL) to model such compositional nature for multi-label image classification. In CPCL, each image is represented by a collection of category-related features used to eliminate the negative effects from location information. Then, a compositional learning module is introduced to compose and decompose the category-related features with their corresponding category prototypes, which are derived from the semantic representations of categories. If the image has the given object, the output after composing should be closer to the original input than the output after decomposing. Contrarily, if the image does not have the given object, the output after decomposing should be closer to the original input than the output after composing. We introduce the Transformed Appearance Distance (TAD) to measure the appearance change between the composed and decomposed features relative to the category-related features with respect to each category. Finally, multi-label image classification is accomplished by performing a TAD-based metric learning. Experimental results on three multi-label image classification benchmarks,i.e., NUS-WIDE, MS-COCO and VOC 2007, validate the effectiveness and superiority of our work in comparison with the state-of-the-arts. The source codes of our model have been released onhttps://github.com/ZFT-CQU/CPCL. Fengtao Zhou, Sheng Huang 0001, Bo Liu 0005, Dan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Deep Semantic Dictionary Learning for Multi-label Image ClassificationabstractCompared with single-label image classification, multi-label image classification is more practical and challenging. Some recent studies attempted to leverage the semantic information of categories for improving multi-label image classification performance. However, these semantic-based methods only take semantic information as type of complements for visual representation without further exploitation. In this paper, we present an innovative path towards the solution of the multi-label image classification which considers it as a dictionary learning task. A novel end-to-end model named Deep Semantic Dictionary Learning (DSDL) is designed. In DSDL, an auto-encoder is applied to generate the semantic dictionary from class-level semantics and then such dictionary is utilized for representing the visual features extracted by Convolutional Neural Network (CNN) with label embeddings. The DSDL provides a simple but elegant way to exploit and reconcile the label, semantic and visual spaces simultaneously via conducting the dictionary learning among them. Moreover, inspired by iterative optimization of traditional dictionary learning, we further devise a novel training strategy named Alternately Parameters Update Strategy (APUS) for optimizing DSDL, which alternately optimizes the representation coefficients and the semantic dictionary in forward and backward propagation. Extensive experimental results on three popular benchmarks demonstrate that our method achieves promising performances in comparison with the state-of-the-arts. Our codes and models have been released. Fengtao Zhou, Sheng Huang 0001 |
AAAI | 1 |
| 2021 | Generally Boosting Few-Shot Learning with HandCrafted FeaturesabstractExisting Few-Shot Learning (FSL) methods predominantly focus on developing different types of sophisticated models to extract the transferable prior knowledge for recognizing novel classes, while they almost pay less attention to the feature learning part in FSL which often simply leverage some well-known CNN as the feature learner. However, feature is the core medium for encoding such transferable knowledge. Feature learning is easy to be trapped in the over-fitting particularly in the scarcity of the training data, and thereby degenerates the performances of FSL. The handcrafted features, such as Histogram of Oriented Gradient (HOG) and Local Binary Pattern (LBP), have no requirement on the amount of training data, and used to perform quite well in many small-scale data scenarios, since their extractions involve no learning process, and are mainly based on the empirically observed and summarized prior feature engineering knowledge. In this paper, we intend to develop a general and simple approach for generally boosting FSL via exploiting such prior knowledge in the feature learning phase. To this end, we introduce two novel handcrafted feature regression modules, namely HOG and LBP regression, to the feature learning parts of deep learning-based FSL models. These two modules are separately plugged into the different convolutional layers of backbone based on the characteristics of the corresponding handcrafted features to guide the backbone optimization from different feature granularity, and also ensure that the learned feature can encode the handcrafted feature knowledge which improves the generalization ability of feature and alleviate the over-fitting of the models. Three recent state-of-the-art FSL approaches are leveraged for examining the effectiveness of our method. Extensive experiments on miniImageNet, CIFAR-FS and FC100 datasets show that the performances of all these FSL approaches are well boosted via applying our method on all three datasets. Our codes and models have been released. Yi Zhang 0113, Sheng Huang 0001, Fengtao Zhou |
ACM Multimedia | 3 |
| 2021 | Fusing hypergraph spectral features for shilling attack detection
Hao Li 0137, Min Gao 0001, Fengtao Zhou, Qilin Fan, Yanyan Yang 0002 |
J. Inf. Secur. Appl. | 3 |