VLDB 2026 Research / reviewers in the wild / expert
Lei Fan 0007
dblp:40/759-7
· DBLP profile ↗
38ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0001-9472-7152ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 14 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flexible Concept Bottleneck ModelabstractConcept bottleneck models (CBMs) improve neural network interpretability by introducing an intermediate layer that maps human-understandable concepts to predictions. Recent work has explored the use of vision-language models (VLMs) to automate concept selection and annotation. However, existing VLM-based CBMs typically require full model retraining when new concepts are involved, which limits their adaptability and flexibility in real-world scenarios, especially considering the rapid evolution of vision-language foundation models. To address these issues, we propose Flexible Concept Bottleneck Model (FCBM), which supports dynamic concept adaptation, including complete replacement of the original concept set. Specifically, we design a hypernetwork that generates prediction weights based on concept embeddings, allowing seamless integration of new concepts without retraining the entire model. In addition, we introduce a modified sparsemax module with a learnable temperature parameter that dynamically selects the most relevant concepts, enabling the model to focus on the most informative features. Extensive experiments on five public benchmarks demonstrate that our method achieves accuracy comparable to state-of-the-art baselines with a similar number of effective concepts. Moreover, the model generalizes well to unseen concepts with just a single epoch of fine-tuning, demonstrating its strong adaptability and flexibility. Xingbo Du, Qiantong Dou, Lei Fan 0007 |
AAAI | 3 |
| 2026 | DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation
Yongjia Ma, Lei Fan 0007, Donglin Di, Tonghua Su |
Comput. Vis. Image Underst. | 3 |
| 2026 | OSTE: Omni-Scene Text Editing with Latent Decoupling
Tonghua Su, Fuxiang Yang, Lei Fan 0007, Donglin Di, Zhongjie Wang 0003, Xiangqian Wu 0002 |
Comput. Vis. Image Underst. | 3 |
| 2026 | Slide-aware deep feature prompting for enhanced whole slide image classificationabstractThe advent of Whole Slide Imaging (WSI) has revolutionised digital pathology by enabling computational analysis of gigapixel-scale images. To handle their large size, most deep learning models divide WSIs into patches and apply Multiple Instance Learning (MIL) for slide-level classification. However, MIL models often depend on pre-trained feature extractors, resulting in domain gaps between natural and pathological images. Parameter-Efficient Fine-Tuning (PEFT) via visual prompting has emerged to bridge this gap with minimal overhead. Nevertheless, existing visual prompts are typically attached at the image level and tightly coupled with specific architectures such as CNNs or ViTs, limiting generalisability and scalability in WSI tasks. To overcome these limitations, we propose Slide-aware Deep Feature Prompt (S-DFP), a novel visual prompting method which derives task-specific information directly from feature embeddings and is initialised with slide-specific cues, thereby enhancing compatibility with diverse feature extractors and MIL frameworks. Experiments on four benchmark datasets, CAMELYON16, BRIGHT, TCGA-IDH, and UniToPath, demonstrate that S-DFP consistently boosts MIL model performance by 2–5% in AUC while introducing less than 0.02% additional parameters. Furthermore, when integrated with recent pathology foundation models, S-DFP yields additional performance gains. The code is publicly available at S-DFP . Cong Cong 0001, Yang Song 0001, Antonio Di Ieva, Qiangguo Jin, Lei Fan 0007, Angela Chou, Anthony J. Gill, Sidong Liu |
Expert Syst. Appl. | 5 |
| 2026 | Medical hierarchical image classification via dual-geometry image-text learningabstractHierarchical image classification is a fundamental challenge in medical image analysis, as tree-structured taxonomies inherently reflect biological and clinical relationships, spanning the general categorisation of disease entities and fine-grained cellular distinctions. Existing approaches primarily rely on multi-task learning and fine-grained detection, often requiring intricate model design and complex training strategies. In this paper, we aim to exploit the negative curvature property of hyperbolic space, which allows efficient representation of hierarchical structures. We propose a dual-geometry image-text framework, termed H 2 CL. Specifically, we introduce a lightweight classifier head on top of image backbones to extract both Euclidean and hyperbolic features, which are then combined to simultaneously preserve taxonomic consistency from an etiological perspective and enhance instance discrimination from a morphological perspective. Furthermore, a text branch is incorporated to integrate label semantics, where an entailment loss is employed to jointly model image–text alignment and inter-sample relationships. Extensive experiments on cervical cell, skin lesion, and gallbladder disease datasets demonstrate that our framework consistently outperforms advanced methods. Compared to the standard Swin Transformer, H 2 CL achieves an average accuracy improvement of 7% across all three datasets at the fine-grained level, with similarly consistent gains observed when integrated with other backbone models. The source code is publicly available at https://github.com/MCPathology/H2CL . Lei Fan 0007, Arcot Sowmya, Erik Meijering, ZongYuan Ge, Yang Song 0001 |
Medical Image Anal. | 1 |
| 2026 | M 3 Surv : Fusing Multi-slide and Multi-omics for Memory-augmented robust Survival predictionabstractMultimodal survival prediction is crucial for personalized oncology. However, existing methods typically integrate only Formalin-Fixed Paraffin-Embedded (FFPE) slides with a single omics type, such as genomics, overlooking Fresh Frozen (FF) slides that better preserve molecular information, as well as richer multi-omics data like proteomics and transcriptomics. More critically, the complete absence of certain modalities due to clinical constraints ( e.g. , time or cost) severely limits the applicability of conventional fusion models that rely on inter-modality correlations. To address these gaps, we propose M 3 Surv, a framework designed to integrate multi-pathology slides (both FF and FFPE) with multi-omics profiles. For multi-slide fusion, we design a divide-and-conquer hypergraph learning approach to capture both intra-slide higher-order cellular structures and inter-slide relationships, yielding a unified pathology representation. To enrich the biological context, we integrate multi-omics data and employ interactive cross-attention to fuse the pathological and omics modalities. To tackle the missing modality, we introduce a prototype-based memory bank. During training, this memory bank learns and stores representative pathology-omics feature prototypes. At inference, even if a modality is entirely missing, the model can query the bank with available features and robustly impute information from the most similar prototype. Extensive experiments on five TCGA cancer datasets and an in-house dataset demonstrate that M 3 Surv outperforms state-of-the-art methods, achieving an average 2.2% improvement in C-Index. The framework also shows strong stability across various missing modality scenarios, highlighting its clinical potential in real-world, data-incomplete scenarios. Mingcheng Qu, Donglin Di, Yue Gao 0002, Yang Song 0001, Lei Fan 0007 |
Medical Image Anal. | 6 |
| 2026 | STAG: Biologically guided spatial transcriptomics prediction via hypergraph learningabstractSpatial transcriptomics (ST) enables spatially resolved gene expression profiling within intact tissue sections. However, its widespread adoption is constrained by the high cost and low throughput of current sequencing-based protocols. This has motivated growing interest in computationally predicting gene expression directly from routinely acquired histology images. Existing methods are largely restricted to isolated 2D tissue slices and fail to capture richer spatial relationships or structured dependencies among spot-level gene expression profiles. In this paper, we propose STAG, a dual-branch framework for gene-aware expression prediction and spatial context modeling. A Query branch predicts ST expression for an individual target spot, while a Neighbor branch acts as an auxiliary branch to model structured relationships among multiple spots. By leveraging hypergraph learning, the Neighbor branch captures higher-order spatial and molecular dependencies, enabling unified modeling of both intra-slice and inter-slice relationships. This design supports standard 2D settings (a single slice) and naturally extends to 3D scenarios when adjacent tissue sections are available. Moreover, STAG leverages gene semantic information as biological guidance by encoding gene names with a foundation model, enabling coordinated gene-aware interactions beyond independent gene prediction. STAG achieves an average gain of 5.16% in PCC@250 across six datasets. Under highly variable gene selection, STAG maintains the lowest RMSE and highest PCC@50 across three datasets. The effectiveness of the learned representations is further demonstrated in pseudo-3D prediction and downstream cancer classification tasks. Code is available at https://github.com/MCPathology/STAG. Mingcheng Qu, Yuchuan Zhao, Donglin Di, Xiu Su, Hongyan Xu 0002, Yang Song 0001, Lei Fan 0007 |
Medical Image Anal. | 8 |
| 2026 | Learning priority-aware controllable poster layout generation
Fuxiang Yang, Wendi Hou, Lei Fan 0007, Tonghua Su, Lingxiao He, Chengzhou Li, Meng Wang 0001, Qianlong Xie, Donglin Di, Xun Yang 0001 |
Pattern Recognit. | 3 |
| 2026 | Noise-aware cross attention for image manipulation localizationabstract• A Gated Noise Extractor that dynamically captures noise features from multiple strategies. • Dual-granularity contrastive learning for more discriminative noise extraction. • Noise-domain guided fusion module t • reduce interference from irrelevant in- formation. • An efficient model with low parameter count and computational complexity. Modern image manipulation techniques have achieved visual realism that often deceives the human eye and semantic-based detectors. However, manipulation operations typically disturb the intrinsic statistical properties of images. Unlike high-level semantic content, which remains visually consistent, such disturbances manifest as anomalies in noise characteristics, including inconsistencies in sensor pattern noise, distinct high-frequency residuals, and unnatural frequency-domain artifacts introduced by resampling or synthesis. These subtle forensic cues provide more reliable evidence for manipulation localization but are often suppressed by standard RGB-domain feature extractors. Existing IML methods often rely on a single noise feature extraction strategy or treat all tampering techniques uniformly, leading to two major limitations, incomplete noise characterization and insufficient tampering-type awareness . We propose a Noise-aware Contrastive localization Network (NC-Net), which introduces two key modules. Firstly, a Gated Noise Extractor that captures mixed noise-domain patterns using a gated network combining features derived from BayarConv and Discrete Wavelet Transform (DWT) operations. This extractor is further enhanced by a dual-granularity contrastive learning strategy, which models distributional discrepancies both within images (between manipulated and authentic regions) and across images (among different manipulation types). Secondly, a Multi-Scale Fusion Module that adaptively integrates noise-domain and RGB-domain semantic features via a cross-domain attention mechanism and a top-down feature pyramid. A lightweight decoder then produces the final localization map with high precision. NC-Net enables end-to-end joint optimization of the noise extraction and RGB branches, achieving state-of-the-art performance with competitive computational overhead. Extensive experiments demonstrate its superiority over existing methods. Source code is available at https://github.com/HIT-liar/NC-Net . Hongshi Zhang, Tonghua Su, Fuxiang Yang, Donglin Di, Yang Song 0001, Lei Fan 0007 |
Pattern Recognit. | 7 |
| 2026 | Energy-Efficient Federated Learning With Dynamic Model Pruning for Industrial IoTabstractWith the advent of the Industry 4.0 era, Federated Learning (FL) provides robust data privacy protection for smart manufacturing and supply chain optimization, while facilitating collaborative intelligent optimization across enterprises and devices. However, the complex and overparameterized deep neural networks used in FL result in significant computational overhead for Industrial Internet of Things (IIoT) devices, leading to low energy efficiency and hindering the practical deployment of FL on IIoT devices. Moreover, the widespread data and device heterogeneity in the IIoT exacerbates the decrease in energy efficiency caused by inconsistent computational efficiency across nodes. This article proposes an energy-efficient dynamic model pruning method for FL, named EDPrune-FL, to address the aforementioned challenges. Compared to existing methods, this approach offers greater flexibility and efficiency by utilizing a dynamic pruning rate allocation mechanism. This mechanism updates the pruning rate for each participating client in every communication round, allowing the pruning upper bound to adapt to the varying importance of different learning stages in FL. EDPrune-FL ensures the global model’s performance while reducing the training energy consumption of clients in heterogeneous environments. To guarantee that dynamic pruning maintains the stability and effectiveness of the model in heterogeneous environments, we also demonstrated the convergence of EDPrune-FL and discussed the relationship between pruning rates and convergence, providing a qualitative analysis. Experimental results demonstrate that our method outperforms the state-of-the-art technique across four real-world datasets. With tests conducted on 100 clients, our approach reduces energy consumption by 10% while maintaining comparable accuracy. Guangsheng Chen, Fangyu Sun, Weitao Zou, Chao Li 0066, Yipeng Zhou, Moule Lin, Peng Liu 0023, Linkang Geng, Lei Fan 0007, Weipeng Jing 0001 |
IEEE Trans. Ind. Informatics | 9 |
| 2026 | Tuning-Free Long Video Generation via Global-Local Collaborative DiffusionabstractCreating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotemporal inconsistencies and high computational resource demands. We propose Global-Local Collaborative Diffusion (GLC-Diffusion), a tuning-free method for long video generation. It models the long video denoising process by establishing denoising trajectories through Global-Local Collaborative Denoising (GLCD) to ensure overall content consistency and temporal coherence between frames. Additionally, we introduce a Noise Reinitialization strategy which combines local noise shuffling with frequency fusion to improve global content consistency and visual diversity. Further, we propose a Video Motion Consistency Refinement (VMCR) module that computes the gradient of pixel-wise and frequency-wise losses to enhance visual consistency and temporal smoothness. Extensive experiments, including quantitative and qualitative evaluations on videos of varying lengths (e.g., 3× and 6× longer), demonstrate that our method effectively integrates with existing video diffusion models, producing coherent, high-fidelity long videos superior to previous approaches. Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie 0009, Lei Fan 0007, Wei Chen 0089, Na Zhao 0004, Xun Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | GRPose: Learning Graph Relations for Human Image Generation with Pose PriorsabstractRecent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this paper, we propose a framework that delves into the graph relations of pose priors to provide control information for human image generation. The main idea is to establish a graph topological structure between the pose priors and latent representation of diffusion models to capture the intrinsic associations between different pose parts. A Progressive Graph Integrator (PGI) is designed to learn the spatial relationships of the pose priors with the graph structure, adopting a hierarchical strategy within an Adapter to gradually propagate information across different pose parts. Besides, a pose perception loss is introduced based on a pretrained pose estimation network to minimize the pose differences. Extensive qualitative and quantitative experiments conducted on the Human-Art and LAION-Human datasets clearly demonstrate that our model can achieve significant performance improvement over the latest benchmark models. Xiangchen Yin, Donglin Di, Lei Fan 0007, Hao Li 0030, Wei Chen 0089, Gouxiao Fei, Yang Song 0001, Xiao Sun 0003, Xun Yang 0001 |
AAAI | 3 |
| 2025 | Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation LearningabstractUniversal, transferable whole-slide image (WSI) representations are central to computational pathology. Incorporating multiple markers (e.g., immunohistochemistry, IHC) alongside H&E enriches H&E-based features with diverse, biologically meaningful information. However, progress is limited by the scarcity of well-aligned multi-stain datasets. Inter-stain Misalignment shifts corresponding tissue across slides, hindering consistent patch-level features and degrading slide-level embeddings. To address this, we curated a slide-level aligned, five-stain dataset (H&E, HER2, KI67, ER, PGR) to enable paired H&E-IHC learning and robust cross-stain representation. Leveraging this dataset, we propose Cross-Stain Contrastive Learning (CSCL), a two-stage pretraining framework: a lightweight adapter trained with patch-wise contrastive alignment to improve the compatibility of H&E features with corresponding IHC-derived contextual cues; and slide-level representation learning with Multiple Instance Learning (MIL), which uses a cross-stain attention fusion module to integrate stain-specific patch features and a crossstain global alignment module to enforce consistency among slide-level embeddings across different stains. Experiments on cancer subtype classification, IHC biomarker status classification, and survival prediction, show consistent gains by yielding high-quality, transferable H&E slide-level representations. The code and data are available at: https://github.com/lily-zyz/CSCL. Yizhi Zhang, Lei Fan 0007, Zhulin Tao, Donglin Di, Yang Song 0001, Sidong Liu, Cong Cong 0001 |
BIBM | 2 |
| 2025 | MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny ObjectsabstractWe present MANTA, a visual-text anomaly detection dataset for tiny objects. The visual component comprises over 137.3K images across 38 object categories spanning five typical domains, of which 8.6K images are labeled as anomalous with pixel-level annotations. Each image is captured from five distinct viewpoints to ensure comprehensive object coverage. The text component consists of two subsets: Declarative Knowledge, including 875 words that describe common anomalies across various domains and specific categories, with detailed explanations for ⟨what, why, how⟩, including causes and visual characteristics; and Constructivist Learning, providing 2K multiple-choice questions with varying levels of difficulty, each paired with images and corresponded answer explanations. We also propose a baseline for visual-text tasks and conduct extensive benchmarking experiments to evaluate advanced methods across different settings, highlighting the challenges and efficacy of our dataset. Lei Fan 0007, Dongdong Fan, Zhiguang Hu, Yiwen Ding, Donglin Di, Kai Yi, Maurice Pagnucco, Yang Song 0001 |
CVPR | 1 |
| 2025 | Prototype-Based Image Prompting for Weakly Supervised Histopathological Image SegmentationabstractWeakly supervised image segmentation with image-level labels has drawn attention due to the high cost of pixel-level annotations. Traditional methods using Class Activation Maps (CAMs) often highlight only the most discriminative regions, leading to incomplete masks. Recent approaches that introduce textual information struggle with histopathological images due to inter-class homogeneity and intra-class heterogeneity. In this paper, we propose a prototype-based image prompting framework for histopathological image segmentation. It constructs an image bank from the training set using clustering, extracting multiple prototype features per class to capture intra-class heterogeneity. By designing a matching loss between input features and class-specific prototypes using contrastive learning, our method addresses inter-class homogeneity and guides the model to generate more accurate CAMs. Experiments on four datasets (LUAD-HistoSeg, BCSS-WSSS, GCSS, and BCSS) show that our method outperforms existing weakly supervised segmentation approaches, setting new benchmarks in histopathological image segmentation.1 Qingchen Tang, Lei Fan 0007, Maurice Pagnucco, Yang Song 0001 |
CVPR | 2 |
| 2025 | Interpretable Image Classification via Non-parametric Part Prototype LearningabstractClassifying images with an interpretable decision-making process is a long-standing problem in computer vision. In recent years, Prototypical Part Networks has gained traction as an approach for self-explainable neural networks, due to their ability to mimic human visual reasoning by providing explanations based on prototypical object parts. However, the quality of the explanations generated by these methods leaves room for improvement, as the prototypes usually focus on repetitive and redundant concepts. Leveraging recent advances in prototype learning, we present a framework for part-based interpretable image classification that learns a set of semantically distinctive object parts for each class, and provides diverse and comprehensive explanations. The core of our method is to learn the partprototypes in a non-parametric fashion, through clustering deep features extracted from foundation vision models that encode robust semantic information. To quantitatively evaluate the quality of explanations provided by ProtoPNets, we introduce Distinctiveness Score and Comprehensiveness Score. Through evaluation on CUB-200-2011, Stanford Cars and Stanford Dogs datasets, we show that our framework compares favourably against existing ProtoPNets while achieving better interpretability. Code is available at: https://github.com/zijizhu/protonon-param. Zhijie Zhu, Lei Fan 0007, Maurice Pagnucco, Yang Song 0001 |
CVPR | 2 |
| 2025 | DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
Donglin Di, Wenzhang Sun, Yongjia Ma, Hao Li 0030, Wei Chen 0089, Lei Fan 0007, Tonghua Su, Xun Yang 0001 |
ICCV | 7 |
| 2025 | Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly Detection
Lei Fan 0007, Donglin Di, Anyang Su, Tianyou Song, Maurice Pagnucco, Yang Song 0001 |
ICCV | 1 |
| 2025 | EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer ModelsabstractDiffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models often focus on more common categories. In large-scale fine-grained image generation, issues of semantic information entanglement and insufficient detail in the generated images still persist. This paper attempts to introduce a concept of a "tiered embedder" in fine-grained image generation, which integrates semantic information from both super and child classes, allowing the diffusion model to better incorporate semantic information and address the issue of semantic entanglement. To address the issue of insufficient detail in fine-grained images, we introduce the concept of super-resolution during the perceptual information generation stage, enhancing the detailed features of fine-grained images through enhancement and degradation models. Furthermore, we propose an efficient ProAttention mechanism that can be effectively implemented in the diffusion model. We evaluate our method through extensive experiments on public benchmarks, demonstrating that our approach outperforms other state-of-the-art fine-tuning methods in terms of performance. Donglin Di, Tonghua Su, Lei Fan 0007 |
ICME | 4 |
| 2025 | Global-Local Aware Scene Text EditingabstractScene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer from two major challenges: inconsistency and length-insensitivity. They often fail to maintain coherence between the edited local patch and the surrounding area, and they struggle to handle significant differences in text length before and after editing. To tackle these challenges, we propose an end-to-end framework called Global-Local Aware Scene Text Editing (GLASTE), which simultaneously incorporates high-level global contextual information along with delicate local features. Specifically, we design a global-local combination structure, joint global and local losses, and enhance text image features to ensure consistency in text style within local patches while maintaining harmony between local and global areas. Additionally, we express the text style as a vector independent of the image size, which can be transferred to target text images of various sizes. We use an affine fusion to fill target text images into the editing patch while maintaining their aspect ratio unchanged. Extensive experiments on real-world datasets validate that our GLASTE model outperforms previous methods in both quantitative metrics and qualitative results and effectively mitigates the two challenges. Fuxiang Yang, Tonghua Su, Donglin Di, Xiangqian Wu 0002, Zhongjie Wang 0003, Lei Fan 0007 |
ICME | 7 |
| 2025 | Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality RebalanceabstractMultimodal pathology-genomic analysis has become increasingly prominent in cancer survival prediction. However, existing studies mainly utilize multi-instance learning to aggregate patch-level features, neglecting the information loss of contextual and hierarchical details within pathology images. Furthermore, the disparity in data granularity and dimensionality between pathology and genomics leads to a significant modality imbalance. The high spatial resolution inherent in pathology data renders it a dominant role while overshadowing genomics in multimodal integration. In this paper, we propose a multimodal survival prediction framework that incorporates hypergraph learning to effectively capture both contextual and hierarchical details from pathology images. Moreover, it employs a modality rebalance mechanism and an interactive alignment fusion strategy to dynamically reweight the contributions of the two modalities, thereby mitigating the pathology-genomics imbalance. Quantitative and qualitative experiments are conducted on five TCGA datasets, demonstrating that our model outperforms advanced methods by over 3.4% in C-Index performance. Code: https://github.com/MCPathology/MRePath. Mingcheng Qu, Donglin Di, Tonghua Su, Yue Gao 0002, Yang Song 0001, Lei Fan 0007 |
IJCAI | 7 |
| 2025 | Spatially Gene Expression Prediction Using Dual-Scale Contrastive Learning
Mingcheng Qu, Yuncong Wu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007 |
MICCAI (15) | 7 |
| 2025 | Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
Mingcheng Qu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007 |
MICCAI (10) | 7 |
| 2025 | Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities
Junze Wang, Lei Fan 0007, Weipeng Jing 0001, Donglin Di, Yang Song 0001, Sidong Liu, Cong Cong 0001 |
MICCAI (11) | 2 |
| 2025 | LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningabstractWhile large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfactual reasoning remains underexplored. To tackle these challenges, we introduce LLaPa, a vision-language model framework designed for multimodal procedural planning. LLaPa generates executable action sequences from textual task descriptions and visual environmental images using vision-language models (VLMs). Furthermore, we enhance LLaPa with two auxiliary modules to improve procedural planning. The first module, the Task-Environment Reranker (TER), leverages task-oriented segmentation to create a task-sensitive feature space, aligning textual descriptions with visual environments and emphasizing critical regions for procedural execution. The second module, the Counterfactual Activities Retriever (CAR), identifies and emphasizes potential counterfactual conditions, enhancing the model's reasoning capability in counterfactual scenarios. Extensive experiments on ActPlan-1K and ALFRED benchmarks demonstrate that LLaPa generates higher-quality plans with superior LCS and correctness, outperforming advanced models. The code and models are available https://github.com/sunshibo1234/LLaPa. Shibo Sun, Xue Li 0011, Donglin Di, Lanshun Nie, Weinan Zhang 0003, Dechen Zhan, Yang Song 0001, Lei Fan 0007 |
ACM Multimedia | 9 |
| 2025 | SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware AlignmentabstractWhile Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle with industrial anomaly detection and reasoning, particularly in delivering interpretable explanations and generalizing to unseen categories. This limitation stems from the inherently domain-specific nature of anomaly detection, which hinders the applicability of existing VLMs in industrial scenarios that require precise, structured, and context-aware analysis. To address these challenges, we propose SAGE, a VLM-based framework that enhances anomaly reasoning through Self-Guided Fact Enhancement (SFE) and Entropy-aware Direct Preference Optimization (E-DPO). SFE integrates domain-specific knowledge into visual reasoning via fact extraction and fusion, while E-DPO aligns model outputs with expert preferences using entropy-aware optimization. Additionally, we introduce AD-PL, a preference-optimized dataset tailored for industrial anomaly reasoning, consisting of 28,415 question-answering instances with expert-ranked responses. To evaluate anomaly reasoning models, we develop Multiscale Logical Evaluation (MLE), a quantitative framework analyzing model logic and consistency. SAGE demonstrates superior performance on industrial anomaly datasets under zero-shot and one-shot settings. The code, model, and dataset are available at https://github.com/amoreZgx1n/SAGE. Guoxin Zang, Xue Li 0011, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song 0001, Lei Fan 0007 |
ACM Multimedia | 7 |
| 2025 | GAOT: Generating Articulated Objects Through Text-Guided Diffusion ModelsabstractArticulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object representations, we propose GAOT, a three-phase framework that generates articulated objects from text prompts, leveraging diffusion models and hypergraph learning in a three-step process. Lei Fan 0007, Donglin Di, Shaohui Liu |
MMAsia | 2 |
| 2025 | Multi-modal hypergraph contrastive learning for medical image segmentation
Weipeng Jing 0001, Junze Wang, Donglin Di, Yang Song 0001, Lei Fan 0007 |
Pattern Recognit. | 6 |
| 2025 | Learning Frequency-Domain Fusion for Multimodal Remote Sensing Semantic Segmentation
Guangsheng Chen, Fangyu Sun, Weipeng Jing 0001, Weitao Zou, Donglin Di, Yang Song 0001, Lei Fan 0007 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | GrainBrain: Multiview Identification and Stratification of Defective Grain KernelsabstractGrain appearance inspection is crucial for evaluating grain quality and determining seed stratification. Typically, trained inspectors manually examine each grain kernel to identify and remove defective ones, which is time-consuming and error-prone. In this article, we present GrainBrain, a robotic vision-based system comprising a hardware prototype (A100) and a deep learning model (GrainAD). A100 is equipped with five cameras to capture high-quality, multiview images of each kernel. The identification of defective kernels is treated as an unsupervised anomaly detection task. GrainAD trains a classifier to distinguish between healthy and pseudoanomaly samples generated at both image and feature levels, and a supervised contrastive learning loss is employed to obtain compact feature representations of healthy kernels. In addition, we release a large-scale dataset containing over 100K annotated images of four types of cereal grains. Extensive experiments were conducted to verify the superiority of our system, achieving an average AUROC of 94.4/90.4% at the image/pixel level. Our system excelled in both efficiency and consistency, as demonstrated by experiments comparing human experts to the system. Lei Fan 0007, Dongdong Fan, Yiwen Ding, Donglin Di, Maurice Pagnucco, Yang Song 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Boundary-Guided Learning for Gene Expression Prediction in Spatial TranscriptomicsabstractSpatial transcriptomics (ST) has emerged as an advanced technology that provides spatial context to gene expression. Recently, deep learning-based methods have shown the capability to predict gene expression from WSI data using ST data. Existing approaches typically extract features from images and the neighboring regions using pretrained models, and then develop methods to fuse this information to generate the final output. However, these methods often fail to account for the cellular structure similarity, cellular density and the interactions within the microenvironment.In this paper, we propose a framework named BG-TRIPLEX, which leverages boundary information extracted from pathological images as guiding features to enhance gene expression prediction from WSIs. Specifically, our model consists of three branches: the spot, in-context and global branches. In the spot and in-context branches, boundary information, including edge and nuclei characteristics, is extracted using pretrained models. These boundary features guide the learning of cellular morphology and the characteristics of microenvironment through Multi-Head Cross-Attention. Finally, these features are integrated with global features to predict the final output.Extensive experiments were conducted on three public ST datasets. The results demonstrate that our BG-TRIPLEX consistently outperforms existing methods in terms of Pearson Correlation Coefficient (PCC). This method highlights the crucial role of boundary features in understanding the complex interactions between WSI and gene expression, offering a promising direction for future research. Codes are available at: https://github.com/WcloudC0416/BG-TRIPLEX Mingcheng Qu, Yuncong Wu, Donglin Di, Anyang Su, Tonghua Su, Yang Song 0001, Lei Fan 0007 |
BIBM | 7 |
| 2023 | Identifying the Defective: Detecting Damaged Grains for Cereal Appearance InspectionabstractCereal grain plays a crucial role in the human diet as a major source of essential nutrients. Grain Appearance Inspection (GAI) serves as an essential process to determine grain quality and facilitate grain circulation and processing. However, GAI is routinely performed manually by inspectors with cumbersome procedures, which poses a significant bottleneck in smart agriculture. In this paper, we endeavor to develop an automated GAI system: AI4GrainInsp. By analyzing the distinctive characteristics of grain kernels, we formulate GAI as a ubiquitous problem: Anomaly Detection (AD), in which healthy and edible kernels are considered normal samples while damaged grains or unknown objects are regarded as anomalies. We further propose an AD model, called AD-GAI, which is trained using only normal samples yet can identify anomalies during inference. Moreover, we customize a prototype device for data acquisition and create a large-scale dataset including 220K high-quality images of wheat and maize kernels. Through extensive experiments, AD-GAI achieves considerable performance in comparison with advanced AD methods, and AI4GrainInsp has highly consistent performance compared to human experts and excels at inspection efficiency over 20× speedup. The dataset, code and models will be released at https://github.com/hellodfan/AI4GrainInsp. Lei Fan 0007, Yiwen Ding, Dongdong Fan, Maurice Pagnucco, Yang Song 0001 |
ECAI | 1 |
| 2023 | Cancer Survival Prediction From Whole Slide Images With Self-Supervised Learning and Slide ConsistencyabstractHistopathological Whole Slide Images (WSIs) at giga-pixel resolution are the gold standard for cancer analysis and prognosis. Due to the scarcity of pixel- or patch-level annotations of WSIs, many existing methods attempt to predict survival outcomes based on a three-stage strategy that includes patch selection, patch-level feature extraction and aggregation. However, the patch features are usually extracted by using truncated models (e.g. ResNet) pretrained on ImageNet without fine-tuning on WSI tasks, and the aggregation stage does not consider the many-to-one relationship between multiple WSIs and the patient. In this paper, we propose a novel survival prediction framework that consists of patch sampling, feature extraction and patient-level survival prediction. Specifically, we employ two kinds of self-supervised learning methods, i.e. colorization and cross-channel, as pretext tasks to train convnet-based models that are tailored for extracting features from WSIs. Then, at the patient-level survival prediction we explicitly aggregate features from multiple WSIs, using consistency and contrastive losses to normalize slide-level features at the patient level. We conduct extensive experiments on three large-scale datasets: TCGA-GBM, TCGA-LUSC and NLST. Experimental results demonstrate the effectiveness of our proposed framework, as it achieves state-of-the-art performance in comparison with previous studies, with concordance index of 0.670, 0.679 and 0.711 on TCGA-GBM, TCGA-LUSC and NLST, respectively. Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2022 | GrainSpace: A Large-scale Dataset for Fine-grained and Domain-adaptive Recognition of Cereal GrainsabstractCereal grains are a vital part of human diets and are important commodities for people's livelihood and international trade. Grain Appearance Inspection (GAI) serves as one of the crucial steps for the determination of grain quality and grain stratification for proper circulation, storage and food processing, etc. GAI is routinely performed manually by qualified inspectors with the aid of some hand tools. Automated GAI has the benefit of greatly assisting inspectors with their jobs but has been limited due to the lack of datasets and clear definitions of the tasks. In this paper we formulate GAI as three ubiquitous computer vision tasks: fine-grained recognition, domain adaptation and out-of-distribution recognition. We present a large-scale and publicly available cereal grains dataset called GrainSpace. Specifically, we construct three types of device prototypes for data acquisition, and a total of 5.25 million images determined by professional inspectors. The grain samples including wheat, maize and rice are collected from five countries and more than 30 regions. We also develop a comprehensive benchmark based on semi-supervised learning and self-supervised learning techniques. To the best of our knowledge, GrainSpace is the first publicly released dataset for cereal grain inspection, https://github.com/hellodfan/GrainSpace. Lei Fan 0007, Yiwen Ding, Dongdong Fan, Donglin Di, Maurice Pagnucco, Yang Song 0001 |
CVPR | 1 |
| 2022 | Fast FF-to-FFPE Whole Slide Image Translation via Laplacian Pyramid and Contrastive Learning
Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001 |
MICCAI (2) | 1 |
| 2021 | Learning Visual Features by Colorization for Slide-Consistent Survival Prediction from Whole Slide Images
Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001 |
MICCAI (8) | 1 |
| 2021 | Enhancing feature fusion with spatial aggregation and channel fusion for semantic segmentationabstractAbstract Semantic segmentation is crucial to the autonomous driving, as an accurate recognition and location of the surrounding scenes can be provided for the street scenes understanding task. Many existing segmentation networks usually fuse high‐level and low‐level features to boost segmentation performance. However, the simple fusion may impose a limited performance improvement because of the gap between high‐level and low‐level features. To alleviate this limitation, we respectively propose spatial aggregation and channel fusion to bridge the gap. Our implementation, inspired by the attention mechanism, consists of two steps: (1) Spatial aggregation relies on the proposed pyramid spatial context aggregation module to capture spatial similarities to enhance the spatial representation of high‐level features, which is more effective for the latter fusion. (2) Channel fusion relies on the proposed attention‐based channel fusion module to weight channel maps on different levels to enhance the fusion. In addition, the complete network with U‐shape structure is constructed. A series of ablation experiments are conducted to demonstrate the effectiveness of our designs, and the network achieves mIoU score of 81.4% on Cityscapes test dataset and 84.6% on PASCALVOC 2012 test dataset. Huifang Kong, Lei Fan 0007 |
IET Comput. Vis. | 3 |
| 2020 | Energy management strategy for electric vehicles based on deep Q-learning using Bayesian optimization
Huifang Kong, Jiapeng Yan, Hai Wang 0004, Lei Fan 0007 |
Neural Comput. Appl. | 4 |